
David Finkelstein
Chakra Team
Summary
Every task we ship has to be difficult, realistic, reliably verifiable, and hard to hack. Two of those can be assessed by running the task, so how fast a trial starts sets how many tasks we can prove are high quality. Our original shared cross-trial infrastructure showed poor tail latencies ; moving them to Daytona on-demand sandboxes cut P99 start time by three orders of magnitude.
At Chakra, quality control is a mix of automated and manual review with models acting as the initial gates and experts reviewing and iterating on the outputs. Every task passes through a substantial review process before it ships, so it is where most of our expert time and compute is spent. If environment development is the factory floor, this process is the test bench at the end of it, ensuring every task is difficult, realistic, reliably verifiable, and hard to hack.
Thorough review is slow, and we are not willing to review less. That leaves one option to increase throughput: make QC faster. This post covers our most recent step - replacing our Kubernetes trial queue with Daytona's on-demand sandboxes.
Two legs to stand on

In our previous post, we discussed our four pillars of good task design. Of the four, difficulty and verifiability can only be checked with rollouts which need to clear internal QC before customer delivery.
Difficult: we run each task at least 10 times against each frontier model and record the pass rate once all runs are complete.
Reliably verifiable: the audit depends on what the verifier is.
For RLVR tasks the verifier is code, and a judge compares every rollout against its verdict to flag false positives and false negatives, following the DeepSWE methodology.
For LLM-judge tasks the verifier is model graded, and we measure judge agreement across solutions sampled along the difficulty curve.
Human reviewers work the same rollouts by hand in both cases. This gives the verifier's error rate and the judge's agreement with a human reviewer. Either kind of error means adjusting the verifier and rerunning the set.
Sampling the difficulty curve takes more rollouts per task than the difficulty check does, and every adjustment means running the set again.
A tail of two schedulers
When optimizing for throughput the batch is the unit that matters because a difficulty check or a verifier rerun is only done when its slowest trial completes. A 1 percent tail on a single trial is a 5 percent tail on a batch of five (1 - 0.99⁵), so one batch in twenty waits on it.
Before partnering with Daytona, an SQS dispatcher queued each trial onto a Kubernetes cluster, which spun up a single-use pod per trial. Start time depended on pod scheduling and cluster contention rather than on the trial itself, and the tail of the distribution was measured in hours. Optimizing with warm pools was viable, but came with the cost of infrastructure maintenance taking time and resources we prefer to spend on research with our customers and experts.
Instead, since Q1 we've been migrating trials onto Daytona sandboxes in stages; each trial now requests an isolated sandbox directly. Two reasons made Daytona the fit:
Harbor framework has first class support for Daytona and the thoughtful integration surface allows us to maximize parallelism.
Daytona schedules sandboxes directly on bare metal rather than through cluster orchestration, which allows for much faster cold-start times.
We compared 58,596 Kubernetes trials from October 2025 through May 2026 against 1,135 Daytona trials sampled in August 2026.

Start latency | Kubernetes | Daytona |
p50 | 158 s | 3.5 s |
p90 | 2.8 h | 6.1 s |
p95 | 5.7 h | 23 s |
p99 | 16.7 h | 46 s |
max | 22.1 h | 287 s |
A p99 that moves from over 16 hours to 46 seconds is a large enough change to alter real team workflows:
Experts can now adjust a verifier, rerun the set, and read the result in one sitting.
Verifier audits run on the full trial set.
Sandboxes come up fast enough that the model provider's rate limit is now what caps how many trials run at once.
Cluster sizing no longer requires manual management.
On our busiest day we started 7,544 trials, a 51 percent increase over the old peak, and on a list-price basis infrastructure cost per trial fell roughly 85 percent.
RL capacity is difficult to plan because the workload looks like a square wave. You go from flat to full utilization for hours at a time, then back down. For evals, each trial needs a clean sandbox so one run can't contaminate the next. When you kill one, you want the next one back instantly. Chakra’s migration is a real-world example of how Daytona removes that provisioning bottleneck for teams running RL and evals at scale without overbuying capacity.
- Ivan Burazin, Co-Founder & CEO at Daytona
With this new infrastructure in place, the model provider's rate limit is the new constraint to work on.
What gets measured gets upgraded
At Chakra, quality is a non-negotiable fixed cost. Every task embodies the same four pillars regardless of demand, so when demand stretched our capabilities the only variables left to focus on were time and compute.
We went after those the way we go after a task: measure it, find where it fails, change one thing, measure again. And to the extent we are presented with the opportunity to work with and learn from a team that holds their layer of the stack to the same standard, we take it every time.
With that, in close, a very special thank you to the Daytona team for building the infrastructure which made our migration possible. They scaled capacity as fast as we could fill it and treated every setback as a hill to be climbed. We couldn't ask for a better partner in our mission critical endeavors.
As always, we will continue to obsess over task design at Chakra Labs. If you're interested in pushing the edges of model capacity, please reach out.

David Finkelstein
Chakra Team
David is Creative Director at Chakra Labs, where he leads creative and editorial strategy. He began his career in finance before moving into brand strategy, where he's spent over a decade translating complex ideas into clear stories across CPG, wellness, and emerging technology. He studied economics and East Asian Languages and Cultures at the University of Illinois Urbana-Champaign.
