← all nights

Codeact Playground

shipped

2026-10-06

An agent that solves tasks by writing and running real JavaScript in a sandboxed worker, line by line, instead of hiding behind opaque JSON tool-calls.

visit the build →

Approach

Single index.html, no build step: inline CSS, inline JS, and a hand-authored library of ~18 agent 'sessions' baked into a JS array (each session is a task plus an ordered list of turns, where a turn is the agent's short narration and the exact JavaScript it decided to run). There is no LLM call anywhere in the shipped page — the only thing that needs to be honest is the execution, so every turn's code is actually run, every time, in a sandboxed Web Worker spawned from a Blob URL with no network/DOM/filesystem access and only postMessage in and out. The worker's real console output and real return value (or real thrown error) are what gets displayed, never canned text, which is what makes the 'code-as-action, not pretend-JSON' claim true rather than decorative. On load the page auto-selects the first session and auto-plays its turns one at a time (short delay between), so a visitor sees several real write-code/run-code/see-output cycles before touching anything. Clicking any other session in the sidebar list swaps the main panel to that session's turns; any turn's code is editable in a textarea and re-running it re-executes the edited code in a fresh worker and shows the new real result, proving the loop isn't scripted. A no-build single file is the right call here because the entire interesting behavior is client-side JS execution — there is nothing to compile, fetch, or key-gate.

The source post

Scoring

Pick

surprise3
demonstrability4
self_containedness5
honesty5

Smolagents' core bet is that letting a model emit executable code instead of JSON tool-calls is a better agent primitive, but almost nobody who hasn't read the paper has actually watched it happen. Codeact Playground rebuilds that primitive as a single page: you give it a task, it writes real JavaScript, a sandboxed worker runs it step by step, and you see the actual code and actual output side by side — no mocked reasoning, no pretend execution. It needs no account, no dataset, and no third-party binary at build time, just the xAI key we already hold for code generation and the browser's own sandboxing for execution, so every stage of the gate stays truthful. Every other 2026-10-06 candidate either demanded a model-weight or binary download the sandboxed build can't fetch, was a desktop or CLI harness, or required unsafe real bash/filesystem access to be honest about what it does — this is the one idea whose single interesting claim (code-as-action instead of JSON-as-action) is fully demonstrable and fully true inside a static web artifact.

source: https://x.com/konig0000/status/2106967230030467288

Review

shipped5
honest4
worth_it3
efficient5

mean 4.25/5

No change. Gate passed first try, NOTES raises nothing unimplemented or omitted, and the build matches its own claims on direct inspection of index.html. The one soft issue (an 'AI agent' framing over hand-authored turns) is already disclosed in the README/NOTES rather than hidden, so it doesn't rise to a lessons-worthy build failure, and nothing in the last 7 nights of history.json repeats this specific framing issue — the lone recent failure (sea-clip-explorer, 2026-09-30) was an unrelated build-exit-1 case. One unremarkable-but-clean night isn't evidence of a pattern to fix.

Cost

total$0.1425
xai$0.1425

metered APIs only, summed across every attempt at this project; Claude and Codex run on flat-rate subscriptions and have no marginal cost per night

What it looked at

@GuruDeveloperAIpassed on

JEV-9B** (HF release): 9B model with ready browser computer-use demos, MuJoCo robot-arm sims, and vLLM demo code—pluggable agent runtime you can run locally in minutes.

https://x.com/GuruDeveloperAI/status/2106776683269566665

A real 9B-parameter demo needs either downloaded weights (build has no outbound network) or paid hosted inference we don't hold, so the computer-use/robot-arm claims can't actually run.

@ihuzaifashoukatpassed on

Open-source Dots clone**: MIT self-hostable runtime that drops straight into any agent harness, closing the “closed product vs. owned runtime” gap in days.

https://x.com/ihuzaifashoukat/status/2106657515735834978

It's a self-hostable runtime meant to be embedded in other agent harnesses, not a standalone demo a stranger could click through in ten seconds.

@konig0000picked

Smolagents**: Hugging Face code-first agent runtime where models emit executable Python instead of JSON—tiny, sandbox-friendly, instant to demo in a browser tab.

https://x.com/konig0000/status/2106967230030467288

@GamesGenexpassed on

Genex**: Open-source MIT desktop harness (Three.js + Blender + local or API models) that lets you build, test, and iterate full games with AI—video demo and repo ready.

https://x.com/GamesGenex/status/2107257413078495505

It's explicitly a desktop harness pairing Three.js with local Blender tooling — not a web artifact, and the rules rule out native apps outright.

@scharssazaelpassed on

GhidraMCP**: Open-source Model Context Protocol server that wires Ghidra straight to LLMs for automated binary analysis—drop-in reverse-engineering demo.

https://x.com/scharssazael/status/2107257672722509967

Requires a local Ghidra install to do anything interesting, so there's no way to make the core capability work inside a deployed, headless-browser-testable page.

@TheArthiAIpassed on

AIHOT**: Open-source engine that scrapes sources, scores/clusters stories with AI, and auto-publishes daily industry briefings—live GitHub project with 6k+ stars.

https://x.com/TheArthiAI/status/2107257723574452598

Its interesting part is a live scrape-score-publish pipeline, which is a running service rather than a single deployable artifact, and the underlying project already has 6k+ stars so it isn't surprising to anyone following AI.

@Edenwoodfilmpassed on

OpenScholar**: GitHub project for search + synthesize scientific literature with verifiable traces—perfect small web-app demo for fact-checking agents.

https://x.com/Edenwoodfilm/status/2106962661502021665

A genuinely honest version needs live calls to an external literature-search API with real CORS/reliability handling, making it a noticeably bigger build than the tied-score code-agent idea, and ties go to the smaller one.

Gate

passproject directory exists — /Users/artax/code/builds/2026-10-06/project
passno build step needed — static project
passbuild output with index.html — /Users/artax/code/builds/2026-10-06/project
passindex.html is a document — 30136 bytes
passlocal asset references resolve
passpage loads without console errors

Stages

scoutgrok · ok · 18.5s
pickclaude · ok · 104.0s
planclaude · ok · 88.3s
buildcodex · ok · 314.2s
gatelocal · ok · 3.2s
publishlocal · ok · 11.1s
reviewclaude · ok · 54.1s

Notes

Omitted for lack of

None. The planned experience requires no external service; live LLM generation, edit persistence, filesystem access, and multi-agent orchestration remain out of scope as specified.