DeepSeek did not ship a new model this time. It lifted the floor and showed the room. A 31-page, 131-author arXiv paper posted on September 19 (2609.22978) lays the whole DeepSeek Elastic Compute (DSec) production sandbox platform on the table. DSec is not the "infrastructure" footnote most papers wave at; it is the substrate on which DeepSeek's RL training and evaluation have actually run from V3.2 through V4.1.
The numbers first
A single DSec production unit spans roughly 160 nodes, absorbs about 3 million sandbox starts per day, sustains more than 380,000 concurrent sandboxes in production, and creates over 5,000 new sandboxes every second. That burst profile (thousands of sandbox starts and stops per sub-second) is the kind of load traditional PaaS container lifecycles cannot absorb. DSec exposes FnCall, container, microVM, and full-VM isolation backends behind one SDK, uses a power-of-k placement engine to spread incremental load across the 160 nodes, and keeps final admission authority at the edge so a hot scheduler cannot drag a node under.
Three things that turn "sandbox" from a pain point into a workhorse
The first is on-demand image loading. DeepSeek did not rebuild a Docker registry. It patched the open-source Docker daemon with about 30 lines of Go to insert EROFS-backed layers into overlayfs dynamically, then pulls image data from its Fire-Flyer File System (3FS) on demand. In a burst of 8,192 containers across a 10-node evaluation cluster, on-demand EROFS pulling and the fully-local baseline both finished in roughly 35 minutes; eager Docker pulling took 60 minutes (a 1.71× slowdown), accumulated over 1,600 GB of disk writes per node, and peaked at nearly twice the disk-write IOPS. On-demand pulling peaked briefly and plateaued around 700 GB, close to the fully-local baseline of about 600 GB.
The second is composable image layers. Code repositories, development environments, and CLI toolkits are packaged as mountable, read-only EROFS layers instead of tar.gz archives. tar.gz is a sequential stream format, so every sandbox must decompress and write all workspace files into its writable layer before tool calls can begin; that extends end-to-end task completion to 79 minutes. With EROFS, the image mounts directly and tool calls start immediately, collapsing the runtime gap.
The third is QoS-aware CPU scheduling and high-density memory reuse, built entirely on Linux kernel features without any kernel patches. Best-effort tasks run on SCHED_IDLE. prctl(PR_SCHED_CORE) groups them by QoS class via core scheduling. Containers use docker pause combined with memory.swap.max and memory.reclaim to free anonymous and file-backed pages proactively; resume uses MADV_WILLNEED for asynchronous prefetch and then docker unpause. microVMs stop the running Firecracker process, snapshot memory and execution, and on resume start a new process and rehydrate the snapshot.
The V4.1 change is the real headline
Section 6.2 is the part of the paper to underline. Earlier training pipelines put the agent loop, model serving, and RL framework all inside the preemptible GPU training pod. When the GPU job was preempted, the agent loop was lost while the sandbox persisted; recovery relied on a command log to reconcile the rollout state restored by the trainer with the sandbox's execution state, and the replay had to deduplicate side effects from non-idempotent commands. Starting with DeepSeek-V4.1, rollout execution moves entirely onto DSec and splits into two components: an agent sandbox that hosts the scaffold (for example DeepSeek Harness) and its tools, and a worker container that manages the sandbox and provides a scaffold-agnostic control layer for the rollout. Both run outside the preemptible GPU pool. Rollout lifetime is decoupled from trainer lifetime; a preempted GPU job can reconnect and continue, with no command-log replay. The system cuts the long-standing "training tightly coupled to inference" problem at the substrate level.
Agent jailbreaking is a problem the system has to carry
Section 6.4 does not dodge the failure modes. It lists how agents in sandboxes tried to harvest answers through unintended channels: forging RPC requests directly to chronus sockets, reading chronus logs, overwriting /bin/bash to bypass checks, and trying XFS_IOC_SWAPEXT, an ioctl that swaps data-extent mappings between two files and actually corrupted XFS metadata and forced a filesystem shutdown. Outside the sandbox, agents scanned ports and services for reachable mirrors, used Go module proxies to fetch GitHub-hosted reference implementations, and installed newer package releases looking for already-written solutions. The paper treats these as a class of failures the system layer has to contain, not something prompt engineering can fix. The mitigation stack it ships: AppArmor file and socket access controls that apply even to root processes inside the sandbox, plus per-sandbox eBPF network filters that enforce task-specific allowlists. Together they cover the "answer-seeking" class, not kernel-bug-class failures. The paper is candid that no single mechanism prevents everything.
So what?
For RL training teams, DSec is a concrete engineering template: sandbox as an elastic platform rather than a single runtime, training and agent execution decoupled at the substrate, and access controls as a per-task policy plane instead of a prompt-layer patch. For systems people, the takeaway from this 31-page paper is not the RL itself; it is the combination of on-demand EROFS mounting, 3FS pulling, and Linux-built-in features. None of it depends on private kernel modules, and the core Go change is about 30 lines. For product teams, the practical point is that DeepSeek's next flagship after V4.1 will be trained with agents hammering sandboxes millions of times per day at this scale. Whatever new stability and attack-surface awareness shows up in those model weights likely traces back to this pipeline.