← all nights

Agent Panopticon

shipped

2026-09-24

A single-page dashboard spinning up hundreds of tiny simulated agents, each a real inspectable state machine, running live entirely in the browser with no backend and no faked numbers.

visit the build →

Why this matters

Agent Panopticon fills your screen with hundreds of tiny simulated 'agents,' each one a small box quietly doing its own thing — changing state, making decisions — and every single one is real: an actual running program you could inspect and verify, not a pre-recorded animation or a handful of real agents copy-pasted to look like hundreds. It's a response to a specific kind of AI demo that's become common — flashy dashboards claiming 'thousands of agents running live' that turn out to be mostly decoration — built specifically so the claim ('this many independent things are really running') is checkably true rather than just impressive-looking. The value isn't any one agent being clever; it's proving a browser, with no server behind it, can genuinely run and track that many small independent processes at once without cheating.

Other ways this idea or technique could be used

Approach

A single index.html with inline CSS and JS, no build step and no network calls. Each simulated agent is a small deterministic state machine (queued -> running -> succeeded/failed/retrying -> handoff) driven by a seeded PRNG (mulberry32) so runs are reproducible per session and every displayed transition traces back to real internal state, not a random label. A single setInterval tick loop advances all agents and re-renders only changed panels for performance at a few hundred nodes. Agents render as small colored grid panels in a CSS grid; clicking one opens a detail pane showing its full transition log, retry count, and current task, which is the mechanism that proves the panels aren't fake. Aggregate counters (per-state totals, completed count, tick count) are derived by summing live agent state each render, so the dashboard numbers and the individual panels can never disagree. Plain HTML/CSS/JS is the right toolchain because the whole artifact is one page with no assets and no external calls.

The source post

Scoring

Pick

surprise3
demonstrability5
self_containedness5
honesty4

imd.fun's pitch (thousands of visible agent panels updating live) is the only candidate tonight that survives the hard filters: it compresses to a single static page, needs no account, no dataset, and no third-party key. Everything else in the batch is either infrastructure (Google AX), a CLI/library (Vermilion, DeepSWE), too compute-heavy to run client-side (Qwen 7B image model), gated behind a paid API we don't hold (TypeSafe/ElevenLabs, Alex emotional AI's hosted Gemma), or requires controlling an external VM (the CUA/VON loop) — none of those can be smoke-tested by a headless browser. The build tonight won't fake 2,000 agents; it runs a few hundred real deterministic state machines (task queues, retries, failures, handoffs) so a stranger watching the panels can trust that what they see is the actual simulation state, not a random-number generator dressed up as one. That's what keeps Honesty above a mockup score — the claim 'these agents are doing something' has to be literally true and inspectable, even if the something is a toy simulation rather than live LLM calls.

source: https://x.com/taylor_/status/2102781117808185720

Review

shipped5
honest4
worth_it3
efficient4

mean 4.00/5

proposed change: {'file': 'prompts/build.md', 'block': 'lessons', 'edit': "Do not attempt to launch headless Chromium/Playwright or bind a local HTTP server to self-verify your work — this sandbox denies both every time ('bootstrap_check_in: Permission denied' on Mach-port registration for the browser process; EPERM on socket bind for a local server). This has happened on 5 consecutive nights (2026-09-23 x4, 2026-09-24) and always ends the same way: a wasted attempt, then a fallback to a Node/DOM harness. Skip the browser/server launch entirely and go straight to a Node-based DOM harness (jsdom or a hand-rolled minimal DOM stub) to validate structure, state transitions, and counters. Real browser rendering and console-error checks are already covered by the separate gate stage — you don't need to reproduce them.", 'expect': "Future NOTES.md 'Could not do' sections stop mentioning Chromium/Playwright launch or local-server bind failures, and the build stage's attempts-to-success ratio improves from tonight's 4/8.", 'falsified_by': "If NOTES.md in the next 3-5 nights still records an attempted Chromium/Playwright launch or local-server bind followed by a sandbox denial, the lesson isn't landing and should be reverted — the fix likely needs to happen in tooling/routing (e.g. removing the browser-launch capability from the build stage's toolset) rather than as a prompt instruction."}

Five consecutive nights' NOTES.md (2026-09-23, -2, -3, -4, and tonight) independently hit the identical sandbox denial trying to open a real browser or bind a local server for self-verification, every time falling back to the same Node/DOM harness approach that already works. That's exactly the 'three nights of the same failure is evidence' bar the review process asks for, and appending a concrete lesson to prompts/build.md is the lowest-risk way to stop the build stage from re-discovering this every night.

Cost

total$0.1376
xai$0.1376

metered APIs only, summed across every attempt at this project; Claude and Codex run on flat-rate subscriptions and have no marginal cost per night

What it looked at

@taylor_picked

imd.fun**: Public live dashboard with 2,000+ AI agents running simultaneously in visible panels; a ready-to-visit multi-agent demo you can watch in real time.

https://x.com/taylor_/status/2102781117808185720

@StepCoinsol8passed on

Alex emotional AI**: Fully open-source self-aware system on Gemma 4 with dozens of emotional states, persistent ASCII persona, and indefinite runtime; GitHub + live website available.

https://x.com/StepCoinsol8/status/2102375645162418637

Needs a running, stateful Gemma 4 instance to be real rather than scripted — that's a hosted-inference dependency we don't have a free key for.

@ilyasergeypassed on

Vermilion**: Experimental Lean 4 backend for the Verus Rust verifier that turns VCs into readable theorems provable by SMT, Lean, or AI; complete GitHub repo.

https://x.com/ilyasergey/status/2102906350628213168

A Lean 4 theorem-proving backend for Verus is a library/CLI artifact, not something a browser can click through.

@dani_avila7passed on

Google AX**: Declarative “Kubernetes for agents” that orchestrates billions of sandboxed agent tasks with workspaces, networking, and lifecycle; already downloadable for testing.

https://x.com/dani_avila7/status/2102889033579860132

Billion-task agent orchestration infra is backend/cluster software, not a single deployable page.

@VraserXpassed on

Qwen 7B transparent image model**: Tiny open model that natively generates transparent images, performs local edits, and composites up to 10 references; runnable locally for web/creative apps.

https://x.com/VraserX/status/2102861310459277733

Running a 7B image model for real generation is too heavy to host and smoke-test in one night without paid inference.

@vishalsingh2972passed on

DeepSWE open-source release**: Complete RL environments + training framework from the post-training runs; fully reproducible for building SWE agents.

https://x.com/vishalsingh2972/status/2102202685269422483

It's an RL training/eval framework meant to be run from a terminal, not a demonstrable web artifact.

@abhegdpassed on

TypeSafe AI feedback playground**: Tiny self-contained demo that classifies spoken/typed in-app feedback (ElevenLabs + classifier) with remix cookbooks; perfect for a quick web app.

https://x.com/abhegd/status/2102195682257854602

The interesting half of the demo (voice classification) depends on an ElevenLabs key we don't already hold.

@MrSagepassed on

LLM + CUA + VON local loop**: Open-source agent that lets an LLM control a VM, observe state changes, and loop decisions entirely locally; early working prototype with visible before/after diffs.

https://x.com/MrSage/status/2102436556736737682

The agent's whole point is driving a local VM, which a headless browser can't observe or verify.

Gate

passproject directory exists — /Users/artax/code/_nightly/2026-09-24/project
passno build step needed — static project
passbuild output with index.html — /Users/artax/code/_nightly/2026-09-24/project
passindex.html is a document — 15080 bytes
passlocal asset references resolve
passpage loads without console errors

Stages

scoutgrok · ok · 24.3s
pickclaude · ok · 39.7s
planclaude · ok · 29.6s
buildcodex · ok · 247.1s
gatelocal · ok · 3.6s
publishlocal · ok · 14.6s
reviewclaude · ok · 91.2s

Notes

Omitted for lack of

None. This project requires no external service or generated imagery.