Skip to main content

View source on GitHub

Demonstrates maintaining a pool of pre-initialized “warm” sandboxes that late-bind to conversations on demand, taking sandbox startup and initialization out of the end user’s wait. This example demonstrates a technique for any situation where expensive setup must finish before an agent can start working, using OpenHands Cloud APIs. It shows that custom images are not the only viable approach for handling initialization that takes more than a few seconds, and it can also be combined with custom images.
When applications have components that run outside the agent control loop and must be available on the system where the agent is running, a custom image is not the only mechanism for packaging these dependencies.Although OpenHands Enterprise has robust support for custom sandbox images, custom images are not yet supported on OpenHands SaaS. Where they are supported, a custom image may still not fully address the situation:
  • The required setup may change over a shorter period than is practical for rebuilding the image.
  • A significant part of the setup may vary so much by task that an image for each variation is impractical.
Even when using custom images in OpenHands Enterprise, some scenarios require additional tasks to be completed on the running sandbox to make it ready for use. If these tasks take more than a few seconds, the Warm Sandbox Pool technique removes that delay from what an end-user waits for by preparing a pool of pre-initialized sandboxes that late-bind to conversations when an end-user begins to interact with the agent.This same approach can be used to install and prepare application services in sandboxes via API calls available in OpenHands SaaS, providing a viable alternative to custom images for your deployment needs.

How It Works

Instead of waiting for sandbox provisioning and initialization every time a user starts a conversation, this approach:
  1. Maintains a pool of pre-initialized sandboxes (e.g., 3 sandboxes)
  2. Pre-installs dependencies (Ruby, gems, application services) during sandbox preparation
  3. Late-binds conversations - when a user needs a sandbox, one is pulled from the pool with no boot or setup wait
  4. Auto-refills - when pool drops below threshold, new sandboxes are automatically provisioned and prepared in the background
This is particularly valuable when:
  • Setup tasks take more than a few seconds (installing Ruby, gems, starting services)
  • You want consistent, fast conversation startup times
  • You need services running and ready before the agent starts working
  • You’re using OpenHands SaaS, where custom images are not yet supported

Architecture

Prerequisites

  • Python 3.10+
  • OpenHands Cloud API key (OH_API_KEY)
  • uv or standard Python environment

Run It

Installation

Then open http://localhost:12000 in your browser. The controller prints a random access code when it starts (look for ACCESS CODE: in its output). Paste it into the page to sign in. Nothing is served without it, and a restart issues a new code.
Heads up: this creates real sandboxes. The controller immediately starts POOL_SIZE sandboxes (default 3) and keeps topping the pool up as you use them. Keep the defaults (pool of 3, threshold 2): a smaller pool leaves nothing to watch, because the refill only starts once ready drops below the threshold. Press Ctrl-C (or send SIGTERM) to stop: every sandbox still sitting in the pool is deleted. Sandboxes already attached to a conversation are left running, like any other conversation sandbox. If the controller is killed with SIGKILL it cannot clean up, so check your sandbox list afterwards.

What to Expect (and what the pool does not speed up)

The pool removes sandbox boot plus your init script from the user’s wait. It does not remove the time OpenHands needs to start the conversation itself. Measured on one beta instance (your numbers will differ, and the web UI shows yours live): So a cold start costs roughly conversation start + warm-up, and a warm start costs conversation start. The more expensive your init script, the bigger the win. For a cheap init script like this demo’s, the saving is modest.

What You’ll See

  1. Initial State: “Preparing pool…” message while the sandboxes initialize (roughly 10-30 seconds each when the platform has capacity, longer otherwise)
  2. Pool Visualization: Real-time status of each sandbox:
    • 🔴 STARTING: OpenHands is provisioning the sandbox
    • 🟡 PREPARING: Installing Ruby, gems, starting service
    • 🟢 READY: Fully initialized and available
  3. Conversation UI: Once pool is ready, type a message to start a conversation
  4. Sandbox Allocation: A ready sandbox is pulled from the pool and attached to your conversation
  5. Auto-Refill: Watch as a new sandbox automatically begins initializing to refill the pool

Configuration

Environment variables:
  • OH_API_KEY: OpenHands Cloud API key (required)
  • OH_API_BASE: API base URL (default: https://app.all-hands.dev). Set this to point at a different OpenHands instance.
  • POOL_SIZE: Target pool size (default: 3)
  • POOL_THRESHOLD: Trigger refill when pool drops below this (default: 2)
  • SANDBOX_SPEC_ID: Optional sandbox spec (runtime image) to use for pool sandboxes
  • INIT_TIMEOUT: Seconds the init script may run in a sandbox (default: 300)
  • MAX_FAILURES: Stop refilling after this many provisioning failures in a row (default: 3). A failed sandbox is deleted immediately, so a broken init script cannot silently create sandboxes forever.
  • HOST: Address the web UI binds to (default: 127.0.0.1). Use 0.0.0.0 to reach it through an OpenHands sandbox work URL. The UI can start conversations with your API key, so it is gated by the access code the controller prints at startup; see the security note above.
  • PORT: Web server port (default: 12000, one of the ports an OpenHands sandbox publishes as a work URL)

Testing the Ruby Service

Once a sandbox is in READY state, the Sinatra service is listening on port 4567 inside the sandbox. That port is not exposed publicly, so reach it from the sandbox itself. The easiest way is to start a conversation and ask the agent:

Demo Application: Ruby Sinatra Service

This example installs a Ruby/Sinatra web service in each sandbox to demonstrate the warm pool technique in a realistic scenario. The Ruby service is only an example: any long-running startup job can be installed and running before agent conversations begin, such as cloning a very large monorepo or downloading the Maven dependencies of a Java monorepo with many projects. The initialization process (installing Ruby runtime, gems, starting the service) shows how to deploy application services using the same API-driven preparation approach for your specific use case.

Implementation Details

Sandbox Initialization Process

The controller uploads sandbox_prep/init_ruby_service.sh and sandbox_prep/quote_service.rb to the sandbox’s agent-server (POST /api/file/upload) and runs the script (POST /api/bash/execute_bash_command). The script does the following (sandboxes run as a non-root user, so it uses sudo for installs):
  1. Install Ruby: apt-get install ruby-full (Ruby 3.3 on current sandbox images)
  2. Install Sinatra: gem install sinatra rackup puma
  3. Deploy Service: Copies the uploaded quote_service.rb into /workspace/services
  4. Start Service: Launches the Sinatra app in the background on port 4567
  5. Verify: Confirms the service responds to health checks
This simulates a realistic scenario where your agent needs specific tools/services pre-installed.

Pool Management

The PoolController class handles:
  • Provisioning: Creates sandboxes via OpenHands Cloud API
  • Monitoring: Polls sandbox status until RUNNING
  • Initialization: Executes preparation scripts via agent-server API
  • Queue Management: Thread-safe queue of ready sandboxes
  • Auto-Refill: Background thread maintains pool size, and stops after MAX_FAILURES consecutive failures
  • Cleanup: Deletes failed sandboxes immediately and all unused pool sandboxes on shutdown
  • Conversation Binding: Attaches conversations to pre-warmed sandboxes

Real-Time Updates

The web UI uses Server-Sent Events (SSE) to stream pool state to the browser every couple of seconds. Besides the sandbox cards it shows live stats (claims, average warm-up, average attach time), the sandboxes already claimed by conversations, and an activity feed of every pool event: created, ready, pulled from the pool, claimed, failed, deleted, refilling halted.

Files

Use Cases

1. Custom Runtime Environments

Pre-install language runtimes (Ruby, Java, Go) that take time to set up:

2. Service Dependencies

Start databases, caches, or mock APIs before the agent runs:

3. Large Codebases

Clone and prepare large repositories with dependencies, such as a very large monorepo or a Java monorepo with heavy Maven dependencies:

4. Custom Application Integration

Replace the demo Sinatra service with your actual application initialization:
By pre-warming sandboxes with your application already running, agents can use your services as soon as the conversation starts, without waiting for sandbox boot and initialization (tens of seconds for this demo, more for heavier setups) on every conversation.

Benefits vs. Custom Images

Warm Sandbox Pool is ideal when:
  • You’re on OpenHands SaaS (custom images not yet supported)
  • Setup time is 10-60 seconds (too slow for UX, too fast to justify custom image complexity)
  • You want flexibility to change initialization without rebuilding images
  • You have predictable conversation volume

Extending the Example

Add More Preparation Steps

Edit sandbox_prep/init_ruby_service.sh to install additional tools:

Customize the Demo Service

Replace sandbox_prep/quote_service.rb with your own Ruby application or gem.

Adjust Pool Parameters

Add Health Checks

Extend the initialization to verify services are actually ready:

Troubleshooting

Pool Never Reaches Ready State

Check the controller’s terminal output and the init log in the web UI. After MAX_FAILURES failures in a row the controller stops creating sandboxes and the UI says so. Common issues:
  • Ruby installation timeout (increase INIT_TIMEOUT)
  • Network issues downloading gems
  • Insufficient sandbox resources

Sandboxes Get Stuck in PREPARING

Failed sandboxes are deleted right away, so to debug the init script keep a sandbox alive and run it by hand. Create one with the start-sandbox/ example (or POST /api/v1/sandboxes), then read its session_api_key and AGENT_SERVER URL from GET /api/v1/sandboxes?id=<sandbox-id>:
Delete the sandbox when you are done (DELETE /api/v1/sandboxes/<id>?sandbox_id=<id>).

High Resource Usage

Reduce POOL_SIZE or implement smarter pool management:
  • Scale pool size based on time of day
  • Implement idle timeout (destroy sandboxes after 30 min unused)
  • Use pool only for peak hours, fall back to JIT otherwise

Next Steps

  1. Production Deployment: Add error handling, logging, metrics
  2. Persistent Storage: Save pool state to Redis/database for crash recovery
  3. Multi-Tenant: Separate pools per user/organization
  4. Dynamic Scaling: Adjust pool size based on demand
  5. Cost Optimization: Implement sandbox recycling (reset instead of destroy)

start-sandbox

Basic sandbox provisioning

clone-and-attach

Conversation attachment patterns

upload-skills

Pre-loading agent skills

License

MIT - See repository root LICENSE file