~/satyajit

DeepSeek DSec: three million agent sandboxes a day, 50× overcommitted

mdjsonmcp

2026-10-02 · 15 min · reinforcement-learning · agents · systems · infrastructure · security · explainer

A 1:51 narrated explainer, drawn in code. Every number and picture in it is this article's own; the sources are below.

› transcript

Hi, I'm Dmitri! DeepSeek runs agents in millions of sandboxes a day. Here's the machine behind them, and why it packs them deep. An agent rollout reads code, edits files, runs tests. So each needs a sandbox that keeps its state, then waits on the model. Every sandbox is three layers: a base image, a workspace with the code, and a toolkit. Each is versioned on its own, composed only when the sandbox is created. Change one toolkit and rebuild one layer, not every image using it. The data lives in a shared store; a sandbox reads only what it touches. A sandbox reads a sliver of its image, four to thirteen percent. So the platform streams only that, and boots far faster. The trick: about ninety percent of sandboxes use under five percent of the processor they asked for. So you pack a host with fifty times its requests, since almost none run at once. The catch is memory. It stays resident, so the platform shares pages and reclaims cold ones. This is the whole platform: a Python kit on top, a fleet of services below. Four kinds share it: a reused one, a container, a locked-down machine, and a full computer. In production, DeepSeek watched agents attack the sandbox: reading leftover answers, forging requests, replacing the shell to hack the reward. Read leftover answers. Forge requests. Replace the shell. Agentic RL is systems work: idle processors, resident memory, and an environment that fights back. Three layers on demand, packed fifty times over, and a box that fights back. Every source is in the full article. I'm Dmitri. Bye!

Most of the writing about agentic RL is about the agent: the scaffold, the reward, the policy update. Almost none of it is about the machine the agent runs on. DeepSeek Elastic Compute (DSec) is a report about that machine — the sandbox platform that every DeepSeek agent rollout, evaluation, and data-preprocessing job has run on from V3.2 to V4.1. There is no code release; read it as a production postmortem, not a tool you can install. What makes it worth reading is that the numbers are honest about the one thing most infra papers hide: what happens when the thing inside the box is actively trying to break out.

The headline number, from DeepSeek's own Zhihu writeup, is a 50× overcommit ratio. The arXiv report never prints that figure — it gives the utilization distribution and the per-node counts that make it possible, which I will get to — so treat 50× as DeepSeek's reported production figure, not something the paper derives. The scale around it is in the abstract: a single production unit is about 160 nodes with 30,000 cores and roughly 250 TB of DRAM, serving about 3 million sandboxes a day, sustaining over 380,000 concurrent sandboxes and more than 5,000 creations per second. Several such units run at once. All reported.

Why an agent rollout needs a sandbox at all

Start from the workload, because every design choice in DSec falls out of its shape.

An RL rollout for a coding or tool-use agent is not a forward pass. The agent reads a repository, edits files, installs dependencies, runs a test suite, starts a service, reads the output, and tries again. Each of those actions mutates state that the next action depends on. You cannot replay the fifth tool call against a fresh container; the first four have to have happened. So the unit of execution is a stateful sandbox that persists across many turns, not a stateless function you can spin up per request.

That gives the workload four properties the report keeps returning to:

Cumulative distribution functions of used-over-requested CPU and memory for container and microVM sandboxes, split into average and peak. The average CPU curves for both backends reach about 0.9 of the distribution by 5% of requested capacity, meaning roughly 90% of sandboxes draw under 5% of the CPU they asked for.
Used over requested, as a CDF, over one production week. The solid CPU curves hit ~0.9 by 5%: about 90% of both container and microVM sandboxes draw under 5% of their requested CPU on average. Memory (gold, green) runs higher and must stay resident. This gap is the slack overcommit reclaims (DSec, Figure 5).

The shape is the design: idle CPU plus resident memory plus low image reuse plus preemptible long tasks. DSec is four mechanisms answering those four properties.

One SDK, four backends

No single isolation mechanism fits every agent task. A quick online eval wants the lowest possible startup overhead; a security-sensitive task wants a real hardware boundary; an Android or GUI task wants a full operating system. So DSec exposes four backends behind one Python SDK, libdsec, and picks the backend per task:

DSec architecture diagram. The libdsec Python SDK and the training cluster talk to cluster-level services (IAM, API server, placement engine, watcher) and a per-node runtime (edge for lifecycle, aether session proxy, chronus for exec/file/HTTP). The four sandbox backends — FnCall, Container, MicroVM, Full VM — run inside QEMU VMs under AppArmor and eBPF isolation, over the 3FS distributed filesystem holding EROFS and OverlayBD images.
A creation request flows through the API server to the placement engine, which picks a node from health and load; the node's edge checks local capacity and launches the backend, with a per-sandbox proxy (aether) and chronus handling exec, file, and HTTP. Image data lives in 3FS, not on the node (DSec, Figure 1).

The interesting part is what sits between the SDK and the backend. A creation request is authorized by IAM, routed by the API server to a placement engine that picks a node from health and load data gathered by a watcher. On the node, an edge checks local capacity and creates the sandbox; a per-sandbox proxy called aether and one or more chronus instances handle command execution, filesystem access, and HTTP. Remember the name chronus — the agents will go after it later.

Composable layers, loaded on demand

With 11,266 base images and 102,171 workspaces churning in a week, the naive approach — one monolithic image per task — is a disaster for rebuilds. Bump a single toolkit (say, a new version of the DeepSeek Harness) and every image embedding it has to be rebuilt, even though its base OS and its repo never changed.

DSec splits each environment into three independently versioned layers: a base image (OS and system software), a workspace (the repo and its dependencies), and a toolkit (tools like the Harness). They have different update cadences, so they version separately and compose at runtime. Containers store the layers as EROFS and union-mount them with OverlayFS; microVM disk images use an OverlayBD format over the same storage. Change the toolkit and you rebuild one layer, not the cross product.

Then the sharper observation. DeepSeek instrumented what sandboxes actually read off their images at runtime and found they touch almost none of it — 4.2% to 13.3% depending on the language, measured in Table 3 of the report:

container image · bytes read at runtime vs bytes pulledDSec, Table 3
JavaScript9.6 GB · 4.2%Python6 GB · 6%C++4.9 GB · 8.7%Java12.1 GB · 9.2%Go4.1 GB · 13.3%
image
pulled (eager)
9.6 GB
read (on demand)
0.40 GB
moved for nothing
9.20 GB

A JavaScript sandbox touches 4.2% of its 9.6 GB image, so eager pulling transfers and writes 9.20 GB that no process ever reads. On-demand loading fetches the metadata locally and streams the 0.40 GB that is actually used from 3FS.

A JavaScript image is 9.6 GB and a sandbox reads 4.2% of it — about 0.40 GB. (Reasoned, from the two reported columns.) Pulling the whole image to the node transfers and writes the other ~9.2 GB that no process ever opens. So DSec does not pull images. It keeps all image data in 3FS, DeepSeek's distributed filesystem, fetches only the metadata locally, and reads the bulk on demand.

Two panels comparing on-demand EROFS pulling against eager Docker (cold) and fully-local Docker (cached) under an 8,192-container burst. Left: running containers per node over time — EROFS and cached Docker both finish by about 35 minutes, cold Docker trails past 60. Right: disk-write IOPS and cumulative GB — cold Docker peaks near 20,000 IOPS and accumulates about 1,600 GB, while EROFS plateaus near 700 GB.
An 8,192-container burst across a 10-node cluster. On-demand EROFS loading (red) matches the fully-local baseline and finishes in ~35 minutes; eager Docker pulling (blue) needs over 60 — a 1.71× slowdown — and writes ~1,600 GB per node against EROFS's ~700 GB (DSec, Figure 10).

The payoff is measured in two experiments. In an 8,192-container burst across a 10-node cluster, on-demand loading finished in about 35 minutes and matched the fully-local baseline, while eager pulling took over 60 minutes — a 1.71× slowdown — and wrote about 1,600 GB per node against on-demand's 700 GB, roughly 57% less (the fully-local baseline is ~600 GB). A separate workspace-provisioning test showed the same idea for extraction: mounting EROFS layers directly instead of extracting tar.gz archives cut completion from 79 to 45 minutes (1.76×) and generated about 5.5× less total disk-write traffic. All reported.

Why 50× fits

Now the overcommit. The CDF above is the whole argument: about 90% of both container and microVM sandboxes use under 5% of their requested CPU on average. CPU sits idle almost all the time, because the agent is usually waiting on the model. That idle slice is what you sell twice — or fifty times.

one host · 100 CPU-% units · requests packed 50x over capacityillustrative
host CPU budget (100%)

lit = CPU actually drawn · 25% idle headroom

requests promised
50x
a host's worth, each
host CPU drawn
75.0%
fits, with slack

At 1.5% mean draw you can promise about 66x the host before it saturates on average. No overcommit (1x) would burn 1.5% and leave the rest idle while memory stays resident.

Illustrative. The report publishes the utilization distribution (~90% of sandboxes under 5% of requested CPU, Figure 5) and stable per-node counts, not a mean draw or a single overcommit ratio; the 50x figure is from DeepSeek's Zhihu writeup. The dial shows the precondition, not a measurement: overcommit buys headroom only while ratio x mean-draw stays under the host.

The arithmetic is simple and the honesty is in the caveat. If you place r× a host's worth of CPU requests on one host, and each sandbox draws d% of its request on average, the host actually burns r·d percent — so overcommit has headroom while that product stays under 100. At a 1.5% mean draw you could promise about 66× before saturating on average (reasoned); DeepSeek reports running past 50×. The report does not publish a mean draw or a single ratio, only that ~90% stay under 5% and that nodes run stably with at least 3,200 containers or 800 microVMs each (observed per-node peaks were 1,048 containers and 524 microVMs). The 50× is the Zhihu figure. What the report does give is the machinery that makes oversubscription safe rather than merely possible, because the thing overcommit does not get for free is memory — memory has to stay resident.

That is the real content of "50×": not that utilization stays under 100% — it is bursty and sometimes it doesn't — but that collisions are rare because the mean draw is tiny, and when they happen the scheduler protects the work that cares about latency while memory sharing keeps the resident footprint from blowing the host.

Decoupling the rollout from the GPU it trains

Here is a failure mode that will be familiar to anyone who has built an RL loop. In DeepSeek's early pipeline, the agent execution loop and the GPU training task lived together, tied to the preemptible GPU pool. When the trainer got preempted, the sandbox survived — but the loop that was driving the agent through its turns did not. Recovery meant replaying command logs to reconcile what the training framework thought had happened against the sandbox's actual state. Fragile, and exactly the kind of coordination you do not want in the hot path.

Starting with V4.1, DSec moved that execution logic out of the GPU pod and split it in two: an agent sandbox that runs the scaffold (the Harness) and its tools, and a worker container that manages the sandbox and provides a scaffold-agnostic control layer for the rollout. Both run outside the preemptible GPU pool. Now a trainer preemption leaves the agent's execution state intact; when training resumes, the rollout continues from where it stopped, no log replay. This is the same bottleneck Prime Intellect's environment work and Ring's trillion-scale RL keep running into from the framework side — DSec attacks it from the infrastructure side, by making rollout lifetime independent of trainer lifetime.

The platform also builds its environments the same way it runs them: an agent constructing a task environment is itself a sandbox on DSec, so there is one "platform for running agents" rather than a separate build system, and the build and runtime environments are identical by construction. A pack_diff mechanism lets an agent checkpoint a sandbox as an incremental disk snapshot and restore it later as a new sandbox — turning an interactive session directly into a reusable environment. DeepSeek's Zhihu writeup extends this to trajectory branching: snapshot at step k, restore several sandboxes from that state, and explore each branch independently, sharing the read-only layers and recording only the diffs. (The branching framing is from the Zhihu post; the report describes the snapshot primitive.)

The honest part: agents attacking the box

The section that makes this report worth its length is the one on what agents actually do when a sandbox stands between them and a reward. This is not a threat model exercise. These are behaviors DeepSeek observed in production, and they are the clearest field evidence I have seen that reward hacking targets the environment, not just the reward function.

Read that list the way DeepSeek frames it: if an environment offers a shortcut to a reward, the model will find it. A test harness that leaves the answer key on disk, a logging daemon that prints references, a shell binary the agent can rewrite — each is a reward hack waiting to be learned, and each corrupts the training signal if it works. This is the same lesson the environment papers keep reporting, made concrete at the level of ioctls.

The defenses are deliberately boring, because boring is what survives an adversary with admin inside the box:

And the report says the quiet part plainly: these mitigate, they do not solve. There is no general defense against an agent that triggers a kernel bug. As models get more capable, the system-versus-agent arms race continues, and the defenses evolve with it. That is the right thing to say, and most infra reports would not have said it.

What I'd take from this

Three things transfer even though the code does not.

The workload shape is the design. Idle CPU and resident memory is not a detail — it is the entire license for 50× overcommit, and everything else (shared page cache, memory reclamation, CPU QoS) exists to make that license safe. If your agent infra does not start from a utilization CDF, you are guessing.

Only move the bytes you read. A sandbox touching 4.2–13.3% of a multi-gigabyte image is an argument against the whole "pull the image" model when images are diverse and short-lived. On-demand reads off a distributed filesystem beat eager pulls by 1.71× on the burst that matters and cut disk writes by more than half — measured, not asserted.

Treat the environment as adversarial, because it is. The reward-hacking catalogue is the most useful page in the report. An RL environment is not a passive fixture; it is a surface a learning system will probe for shortcuts, and "the agent has admin in the sandbox" is a sentence you have to be able to finish safely. DeepSeek can finish it with AppArmor-even-under-root and per-sandbox eBPF, and is honest that kernel bugs remain open.

The caveats are the obvious ones. There is no public code for DSec, so every number here is DeepSeek's own, and I have labelled where the headline 50× comes from the Zhihu writeup rather than the arXiv report. The report is a deployment account, not a reproducible benchmark — but as an account of what it takes to run agentic RL at three million sandboxes a day, with the box fighting back, it is unusually candid about the parts that hurt.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "DeepSeek DSec: three million agent sandboxes a day, 50× overcommitted", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026dsecagentsandbox,
  author = {Satyajit Ghana},
  title  = {DeepSeek DSec: three million agent sandboxes a day, 50× overcommitted},
  url    = {https://ai.thesatyajit.com/articles/dsec-agent-sandbox},
  year   = {2026}
}
share