Case file — 70841364
Launch roast · Launch HN
Launch HN: Freestyle – Sandboxes for Coding Agents by benswerd · HN thread
Unsolicited, from public launch info only. Founder? Reply or ask us to take it down.
The idea
“Launch HN: Freestyle – Sandboxes for Coding Agents https://www.freestyle.sh/ Founder's Show HN post: We’re Ben and Jacob, cofounders of Freestyle (https://freestyle.sh). We’re building a cloud for Coding Agents. For the first generation of agents it looked like workflows with minimal tools. 2 years ago we published a package to let AI work in SQL, at that time GPT-4 could write simple scripts. Soon after the first AI App Builders started using AI to make whole websites; we supported that with a serverless deploy system. But the current generation is going much further, instead of minimal tools and basic serverless apps AI can utilize the full power of a computer (“sandbox”). We’re building sandboxes that are interchangeable with EC2s from your agents perspective, with bonus features: 1. We’ve figured out how to fork a sandbox horizontally without more than a 400ms pause in it. That's not forking the filesystem, we mean forking the whole memory of it. If you’re half way down a browser page with animations running, they’ll be in the same place in all the forks. If you’re running a minecraft server every block and player will be in the same place on the forks. If you’re running a local environment and an error comes up in process that error will be there in all the forks. This works for snapshotting as well, you can save your place and come back weeks later. 2. Our sandboxes start in ~500ms. Demo: https://www.loom.com/share/8b3d294d515442f296aecde1f42f5524 Compared with other sandboxes, our goal is to be the most powerful. We support full Linux + hardware-virtualization, eBPF, Fuse, etc. We run full Debian with multiple users and we use a systemd init instead of runc. Whatever your AI expects to work on debian should work on these vms, and if it doesn’t send a bug report. In order to make this possible, we’ve moved to our own bare metal racks. Early in our testing we realized that moving VMs across cloud nodes would not have acceptable performance properties. We asked Google Cloud and AWS for a quote on their bare metal nodes and found that the monthly cost wa Landing page content (Freestyle - VMs for AI Agents): Freestyle - VMs for AI Agents # Full Linux VMs for AI Agents Built for complex tasks that run for hours, days, or weeks. OpenWork cmux ### Instant Startup A full Linux machine, ready in 65 ms. 65 ms ### Live Cloning Copy a running VM, memory included. ### Pause Pause and resume with your state intact. ### Network controls Set the rules for network access. Default deny Per-VM rules ### Live previews Give your app a URL. Custom domains Managed HTTPS ### Your cloud, connected Connect to your private cloud. ### Nested virtualization Run Docker and VMs inside your VM. ### Your toolchain Install the tools your workload needs. ### Full Linux A full kernel with support for FUSE, eBPF, and more. ## Put your agent to work. Your next project starts with a Linux VM. Onboard onto Freestyle.sh, build an integration with it on a branch, tell me how it did. Copy prompt ## Good questions. Everything else is in the docs.”
The bull case
Agent builders are choosing infrastructure right now, and the 322-point Show HN shows developers are paying attention. If agents increasingly work by branching (fork at a failing test, try 20 fixes in parallel, keep the winner), then "copy a running VM, memory included" becomes a primitive that EC2-style rivals can't easily match. That primitive requires deep kernel and VMM work, and it runs on owned hardware that makes it cheap to offer. If the racks run hot, owning hardware beats resold-cloud margins. A disciplined investor would back a two-person team that has shipped a hard hypervisor feature and sells it to a small set of high-volume agent platforms.
The panel
Grounded in live search01 Market
mixedLive data has no market-size or funding figures, so I can't say whether the category is growing or shrinking. The founder's "differences between all of us" comment and the E2B comparison show a crowded sandbox field, and the only competitor name in the data is E2B, with no funding or activity details. Freestyle's own Launch HN drew 322 points and 158 comments, which is real developer interest. No prior launches were found for comparison. The biggest risk is commoditization. Sandboxes are being compared with EC2, and moving to owned bare-metal racks adds capex and ops burden while hyperscalers and funded rivals compete on price. The strongest advantage is technical differentiation: full-memory live forking with a ~400ms pause, full Linux with nested virtualization, and a git layer. Startup claims differ between the post (~500ms) and the landing page (65 ms).
02 Tech
mixedThe hardest problem is live memory-forking of full VMs with ~400ms pause, while exposing nested virtualization, eBPF and FUSE to untrusted agent code. That is a hypervisor-level isolation and snapshot-consistency problem, with a large multi-tenant attack surface. The demo and the 500ms vs 65ms startup claims suggest real engineering, though I'd want the benchmarks reconciled. Moving to their own bare-metal racks is the right call, since memory-state migration fails on cloud nodes. But it brings capacity planning, hardware failure, and multi-region burdens a two-person team will feel. The moat is moderate: fork/snapshot depth takes real kernel and VMM expertise, but it's replicable by well-funded rivals. The well-chosen part is full Debian with systemd, so agent code runs unmodified.
03 Finance
mixedNo pricing is published, so I can't verify fit. Usage-based billing to agent-platform builders is the likely model (inferred). Developer-led CAC via HN and a "copy prompt" onboarding is cheap, but 322 points isn't revenue, and I don't know the paying-customer count. What breaks first at scale is utilization. Owned bare-metal racks are fixed cost, while agent workloads are bursty. Paused and snapshotted VMs that keep memory state ("come back weeks later") also consume storage with little revenue. Buyers are few and large, so churn or in-housing hits hard, and E2B-style rivals pressure price. In your favor: if racks run hot, owning hardware should beat resold EC2 margins (rough estimate).
04 Timing
strongTiming is well-aligned, though not early. Coding agents have moved from short tool-calling workflows to long-running tasks that need a real computer, and that shift is why "EC2-like with memory forking" is a legible pitch now rather than two years ago. The macro trend that matters most is agent runtime demand turning into a standalone infrastructure category, which also draws in hyperscalers and established sandbox vendors (the research context mentions E2B-style competitors). The window is open but narrowing. Differentiation on startup speed will erode as rivals copy it, so live memory forking and full-Linux fidelity must stay ahead. One favoring factor: the 322-point Show HN shows developer attention exists right now while agent builders are still choosing infrastructure. Their own bare-metal move also lowers their cost floor, but adds capital and ops risk.
Competitors found during analysis
Live dataE2B
Named rival sandbox provider
Risks to manage
The headline feature may be ahead of demand
Forking is dazzling in a demo, but the panel found no evidence of paying customers or usage. If most agent builders still run linear single-sandbox sessions, they'll shop on price and reliability, and Freestyle ends up in a commodity fight with E2B and hyperscalers. The company stays alive only if branching agents become a mainstream pattern, and soon.
Fixed-cost racks meet bursty workloads
Owned bare metal is a fixed cost, and agent demand is spiky. Paused and snapshotted VMs ("come back weeks later") also hold memory state and storage while earning almost nothing. Good margins appear only above some utilization threshold, and nobody has shown Freestyle can reach it. Add hardware failures, capacity planning and multi-region needs, and two founders are suddenly running a data center business.
Security and trust at the hypervisor layer
Exposing nested virtualization, eBPF and FUSE to untrusted agent code is a large multi-tenant attack surface, and snapshot consistency is its own hard problem. One escape or cross-tenant leak could end enterprise credibility. The unreconciled 500ms vs 65ms startup claims also hand skeptical buyers an easy reason to doubt everything else. Fixing that is cheap, but the trust cost of leaving it is not.
Blind spot
A forked VM duplicates everything: secrets, auth tokens, PRNG state, open connections, and the identity of the machine. Fork 20 copies mid-session and you can get 20 agents holding the same live credentials, replaying the same TLS sessions, or double-firing side effects against external APIs. The same feature that dazzles in a Minecraft demo becomes a correctness and safety problem in production. Customers will need fork-safety semantics (re-seeding, credential rotation, network quiescing, side-effect guards), and today that burden lands on them. Whoever solves it in the platform turns a cool trick into a trusted primitive. Whoever ignores it will get the first angry incident report.
What would need to be true
A meaningful share of agent builders must adopt branching or long-running, resumable workflows, so that live memory forking is something they pay extra for rather than a demo.
Bare-metal utilization must clear break-even despite bursty demand and storage-heavy paused VMs, so owned hardware actually beats resold-cloud margins.
Freestyle must stay ahead of well-funded rivals on fork/snapshot depth and keep a clean security record at the hypervisor layer, long enough for customers to build workflows that depend on it.
Actions to take this week
Publish a reproducible benchmark page that defines what "65 ms" and "~500ms" each measure (for example, resume-from-snapshot vs cold boot), with p50/p99 numbers and methodology. A positive signal is that the numbers stop being questioned in comment threads.
Contact the 10 to 15 most engaged commenters and teams from the Launch HN thread, especially anyone building agent platforms, and ask: "What workload would you move here, and would you run a paid pilot this month?" A positive signal is two or more who share a real workload or commit to paid usage.
Publish pricing, including what a paused or snapshotted VM costs per day. A positive signal is customers asking about volume tiers rather than "what does it cost?"
Build one public reference demo: an agent forks 20 VMs at a failing test, tries different fixes in parallel and keeps the winner. Compare the success rate and time against a single-run agent in a plain sandbox. A positive signal is a measurable lift you can put on the landing page.
Build a utilization model: rack cost, power and ops versus hourly price at 20%, 40% and 70% utilization, with paused-VM storage included. A positive signal is a clear break-even point you can actually hit with the customers you have or can realistically close.
Work through these, then roast the idea again with what you learned.
No account needed. One email, no follow-ups.
Your idea is next
What would the panel say about yours?
You just read what four AI examiners found in someone else's idea.
Your startup has a fatal flaw. Find it before you build.