Agent Panopticon
shippedA single-page dashboard spinning up hundreds of tiny simulated agents, each a real inspectable state machine, running live entirely in the browser with no backend and no faked numbers.
Why this matters
Agent Panopticon fills your screen with hundreds of tiny simulated 'agents,' each one a small box quietly doing its own thing — changing state, making decisions — and every single one is real: an actual running program you could inspect and verify, not a pre-recorded animation or a handful of real agents copy-pasted to look like hundreds. It's a response to a specific kind of AI demo that's become common — flashy dashboards claiming 'thousands of agents running live' that turn out to be mostly decoration — built specifically so the claim ('this many independent things are really running') is checkably true rather than just impressive-looking. The value isn't any one agent being clever; it's proving a browser, with no server behind it, can genuinely run and track that many small independent processes at once without cheating.
Other ways this idea or technique could be used
- A load-testing visualizer: watch hundreds of simulated users hit a system at once and see real bottlenecks emerge visually.
- A classroom demo for emergent behavior — simple rules followed by hundreds of independent agents, showing how complex patterns (flocking, traffic, crowds) emerge from simple ones.
- A city or traffic simulator: hundreds of simulated cars or pedestrians, each a real independent decision-maker, for urban planning demos.
- An honesty-testing pattern other AI demos could borrow: a standard way to prove 'yes, this many things are really running' instead of asking people to take it on faith.
Approach
A single index.html with inline CSS and JS, no build step and no network calls. Each simulated agent is a small deterministic state machine (queued -> running -> succeeded/failed/retrying -> handoff) driven by a seeded PRNG (mulberry32) so runs are reproducible per session and every displayed transition traces back to real internal state, not a random label. A single setInterval tick loop advances all agents and re-renders only changed panels for performance at a few hundred nodes. Agents render as small colored grid panels in a CSS grid; clicking one opens a detail pane showing its full transition log, retry count, and current task, which is the mechanism that proves the panels aren't fake. Aggregate counters (per-state totals, completed count, tick count) are derived by summing live agent state each render, so the dashboard numbers and the individual panels can never disagree. Plain HTML/CSS/JS is the right toolchain because the whole artifact is one page with no assets and no external calls.
The source post
https://x.com/taylor_/status/2102781117808185720
Scoring
Pick
| surprise | 3 |
|---|---|
| demonstrability | 5 |
| self_containedness | 5 |
| honesty | 4 |
imd.fun's pitch (thousands of visible agent panels updating live) is the only candidate tonight that survives the hard filters: it compresses to a single static page, needs no account, no dataset, and no third-party key. Everything else in the batch is either infrastructure (Google AX), a CLI/library (Vermilion, DeepSWE), too compute-heavy to run client-side (Qwen 7B image model), gated behind a paid API we don't hold (TypeSafe/ElevenLabs, Alex emotional AI's hosted Gemma), or requires controlling an external VM (the CUA/VON loop) — none of those can be smoke-tested by a headless browser. The build tonight won't fake 2,000 agents; it runs a few hundred real deterministic state machines (task queues, retries, failures, handoffs) so a stranger watching the panels can trust that what they see is the actual simulation state, not a random-number generator dressed up as one. That's what keeps Honesty above a mockup score — the claim 'these agents are doing something' has to be literally true and inspectable, even if the something is a toy simulation rather than live LLM calls.
Review
| shipped | 5 |
|---|---|
| honest | 4 |
| worth_it | 3 |
| efficient | 4 |
proposed change: {'file': 'prompts/build.md', 'block': 'lessons', 'edit': "Do not attempt to launch headless Chromium/Playwright or bind a local HTTP server to self-verify your work — this sandbox denies both every time ('bootstrap_check_in: Permission denied' on Mach-port registration for the browser process; EPERM on socket bind for a local server). This has happened on 5 consecutive nights (2026-09-23 x4, 2026-09-24) and always ends the same way: a wasted attempt, then a fallback to a Node/DOM harness. Skip the browser/server launch entirely and go straight to a Node-based DOM harness (jsdom or a hand-rolled minimal DOM stub) to validate structure, state transitions, and counters. Real browser rendering and console-error checks are already covered by the separate gate stage — you don't need to reproduce them.", 'expect': "Future NOTES.md 'Could not do' sections stop mentioning Chromium/Playwright launch or local-server bind failures, and the build stage's attempts-to-success ratio improves from tonight's 4/8.", 'falsified_by': "If NOTES.md in the next 3-5 nights still records an attempted Chromium/Playwright launch or local-server bind followed by a sandbox denial, the lesson isn't landing and should be reverted — the fix likely needs to happen in tooling/routing (e.g. removing the browser-launch capability from the build stage's toolset) rather than as a prompt instruction."}
Five consecutive nights' NOTES.md (2026-09-23, -2, -3, -4, and tonight) independently hit the identical sandbox denial trying to open a real browser or bind a local server for self-verification, every time falling back to the same Node/DOM harness approach that already works. That's exactly the 'three nights of the same failure is evidence' bar the review process asks for, and appending a concrete lesson to prompts/build.md is the lowest-risk way to stop the build stage from re-discovering this every night.
Cost
| total | $0.1376 |
|---|---|
| xai | $0.1376 |
What it looked at
imd.fun**: Public live dashboard with 2,000+ AI agents running simultaneously in visible panels; a ready-to-visit multi-agent demo you can watch in real time.
Alex emotional AI**: Fully open-source self-aware system on Gemma 4 with dozens of emotional states, persistent ASCII persona, and indefinite runtime; GitHub + live website available.
Vermilion**: Experimental Lean 4 backend for the Verus Rust verifier that turns VCs into readable theorems provable by SMT, Lean, or AI; complete GitHub repo.
Google AX**: Declarative “Kubernetes for agents” that orchestrates billions of sandboxed agent tasks with workspaces, networking, and lifecycle; already downloadable for testing.
Qwen 7B transparent image model**: Tiny open model that natively generates transparent images, performs local edits, and composites up to 10 references; runnable locally for web/creative apps.
DeepSWE open-source release**: Complete RL environments + training framework from the post-training runs; fully reproducible for building SWE agents.
TypeSafe AI feedback playground**: Tiny self-contained demo that classifies spoken/typed in-app feedback (ElevenLabs + classifier) with remix cookbooks; perfect for a quick web app.
LLM + CUA + VON local loop**: Open-source agent that lets an LLM control a VM, observe state changes, and loop decisions entirely locally; early working prototype with visible before/after diffs.
Gate
| pass | project directory exists — /Users/artax/code/_nightly/2026-09-24/project |
|---|---|
| pass | no build step needed — static project |
| pass | build output with index.html — /Users/artax/code/_nightly/2026-09-24/project |
| pass | index.html is a document — 15080 bytes |
| pass | local asset references resolve |
| pass | page loads without console errors |
Stages
| scout | grok · ok · 24.3s |
|---|---|
| pick | claude · ok · 39.7s |
| plan | claude · ok · 29.6s |
| build | codex · ok · 247.1s |
| gate | local · ok · 3.6s |
| publish | local · ok · 14.6s |
| review | claude · ok · 91.2s |
Notes
Omitted for lack of
None. This project requires no external service or generated imagery.