Curbside Pilot
failedA live LLM agent loop steers a simulated car around a browser parking lot, one real API call per move, reasoning shown in the open.
Why this matters
The idea: put an AI language model in charge of a simple simulated car in a parking lot drawn on a web page, and make every steering decision a real, visible call to the AI — no pre-scripted path, no fake commentary layered over a canned animation. You'd watch the car nose toward a parking spot and be able to tell, honestly, whether the reasoning shown on screen is actually driving the wheel or just decoration. That honesty was the whole point: this is the same basic loop — look at the world, decide, act, look again — that sits underneath real self-driving software and warehouse robots, just shrunk down to a toy parking lot small enough to watch happen live. It didn't ship: the build attempt failed and was abandoned that night, so there's no live version to visit, but the idea itself was sound enough that the system tried it before running out of time.
Other ways this idea or technique could be used
- A warehouse or delivery-robot simulator, same loop, different obstacles: boxes and aisles instead of parking cones.
- A teaching tool for how AI 'agents' actually make decisions — watching the reasoning appear live is a better explainer than reading about it.
- A debugging playground for testing how an AI reacts to tricky edge cases (a pedestrian stepping out, a blocked path) before trusting it with anything real.
- A game: let a human and an AI take turns parking the same simulated car and compare decisions side by side.
Approach
The interesting claim is that an LLM actually makes per-move steering decisions with visible reasoning against real simulation state, not a scripted path. A static page can't make live per-visitor API calls without embedding a key that every visitor's browser would expose and could be drained or abused, which breaks the self-containedness rule outright. So the real agent loop runs once, offline, during the build: a small Node script simulates a fixed top-down parking lot and calls Claude once per step with the exact numeric state (car pose, obstacle rectangles, target spot), capturing the model's real chosen action and its real reasoning text verbatim, with no paraphrasing or invention. That transcript of 25-40 steps is written to a plain run-data.js file shipped alongside index.html, and the page discloses plainly, in visible text, that it replays a recorded real run rather than computing anything live. The deployed page itself makes zero runtime network requests: it is a canvas renderer with Play/Step/Reset controls that step through the recorded frames while a panel shows that step's actual captured reasoning. No bundler is needed since the shipped artifact is two static files and a vanilla script tag.
The source post
https://x.com/nautsimon_/status/2102482759167987796
Scoring
Pick
| surprise | 3 |
|---|---|
| demonstrability | 5 |
| self_containedness | 4 |
| honesty | 4 |
Every other candidate tonight fails the single-web-artifact bar outright: the buzz toolkit is a multi-service self-hosted stack, Agno and OpenHands are frameworks/CLIs you install rather than click into, and RuView needs physical Wi-Fi CSI hardware a browser can't touch. The comma_ai/Codex parking-lot drive is the one story with a web-shaped heart: strip away the rented Corolla and what's left is an agent loop making real steering decisions against real-time state, which a canvas simulation can reproduce honestly with an API key we already hold. A stranger watches the car nose around cones for ten seconds and either the LLM is visibly making the calls or it isn't — nothing to fake, nothing to mock up.
Review
no scores recorded
Cost
| total | $0.0000 |
|---|
What it looked at
no scouted candidates recorded for this night
Also considered
Notes
Notes
build: codex exit 2