AI Stem Music Visualizer
shippedUpload or pick a track, AI splits it into stems, and each stem drives its own lighting/physics/camera layer in real time.
Approach
A single index.html with no build step and no external requests. On load, a small in-browser step-sequencer (plain JS, Web Audio API oscillators and noise buffers, no libraries) starts running a pre-authored multi-bar, 32-step pattern across four independent instrument stems — kick, bass, chords, hi-hats — each routed through its own gain node so its envelope can be sampled every animation frame. Visuals are driven from that per-stem envelope data on a requestAnimationFrame loop using four stacked canvas 2D layers (flash-and-shake for the kick, a pulsing low-frequency glow for bass, a shifting background gradient for chords, particle sparkles for hats), plus a live step-grid and per-stem level meters, so the reaction is genuinely tied to which instrument is sounding rather than one broadband FFT bounce. Because AI separation of an arbitrary uploaded recording needs a model or API this environment doesn't have, the page does not accept uploads or claim to isolate real recordings — the four stems are truly separate audio sources by construction, which keeps the 'each stem drives its own layer' claim honest without any ML backend. Visuals start animating immediately from the deterministic pattern clock rather than from AudioContext playback (which browsers may suspend until a user gesture), so the page demonstrates itself with no click required; a real 'Enable Sound' button then starts audio in sync with that same clock. Per-stem mute/solo controls and two alternate preset patterns (including a sparse breakdown section) let a visitor isolate and compare each layer's contribution.
The source post
https://x.com/kaolti/status/2104719243376009287
Scoring
Pick
| surprise | 3 |
|---|---|
| demonstrability | 5 |
| self_containedness | 4 |
| honesty | 4 |
It's the only candidate that's actually a single self-contained web artifact: no account, no dataset, no paid API key, and a stranger can hit play and watch the visuals genuinely react to the drums/bass/vocals rather than a generic FFT bounce. Everything else in the batch fails a hard rule outright — Lasso and OpenMAIC need live source-code access or paid generation APIs, Colibri and Atria Dawn Preview are inference engines/model weights with no web surface (and Atria's 744B params can't run on this box regardless), and Khazix Skills is a folder of agent-skill files, not a deployable artifact. The visualizer is the one idea where the interesting claim (stem-aware visuals) is testable in the artifact itself, not a mockup.
Review
| shipped | 5 |
|---|---|
| honest | 4 |
| worth_it | 3 |
| efficient | 3 |
No change: the pattern of near-max build attempts (5/8, then 7/8, then 7/8 across the last three real build nights) has crossed from one night's weather into a real trend, but run.json only records the final attempt count, not what failed on attempts 1 through 6 -- there is nothing here to turn into a concrete, non-platitude lessons entry, and the 2026-09-28 review already reached this same conclusion for the identical reason. Writing 'reduce build attempts' or similar into prompts/build.md would be exactly the kind of vague instruction the lessons block is supposed to avoid. If per-attempt failure detail becomes available, or if the trend continues past three more nights, that's what would justify a change.
Cost
| total | $0.1388 |
|---|---|
| xai | $0.1388 |
What it looked at
Lasso** (open-sourced dev tool): select any UI element in a running app and have AI directly edit the real source code—perfect for a live, clickable web demo of agentic editing.
Colibri**: zero-dependency C inference engine that runs massive Mixture-of-Experts models locally by spilling across disk/RAM/VRAM—ideal for a tiny self-hosted web playground.
OpenMAIC**: turn one prompt into a complete interactive course (slides, quizzes, videos, AI teachers)—straightforward self-contained web app.
Khazix Skills**: MIT-licensed pack of six reusable, structured agent skills (TDD, cleanup, research, writing, etc.) installable as folders or SKILL.md—drop-in for any coding-agent demo.
Atria Dawn Preview**: 744 B-parameter open-weight coding agent (MIT license, downloadable weights) from Shanghai AI Lab—run the actual model in a local web UI.
AI-stem music visualizer**: browser demo that splits audio into stems via AI then drives visuals (lighting, physics, camera) from drum hits/frequency bands—fully self-contained web experience.
Wanted, and did without
| a hosted audio source-separation model or API (e.g. a Demucs/Spleeter-class endpoint) | letting a visitor upload or pick a real track and see genuinely AI-isolated stems drive the visuals, instead of a procedurally synthesized four-stem demo track |
|---|
Gate
| pass | project directory exists — /Users/artax/code/builds/2026-09-29/project |
|---|---|
| pass | no build step needed — static project |
| pass | build output with index.html — /Users/artax/code/builds/2026-09-29/project |
| pass | index.html is a document — 20866 bytes |
| pass | local asset references resolve |
| pass | page loads without console errors |
Stages
| scout | grok · ok · 26.9s |
|---|---|
| pick | claude · ok · 43.2s |
| plan | claude · ok · 118.7s |
| build | codex · ok · 318.7s |
| gate | local · ok · 3.9s |
| publish | local · ok · 16.4s |
| review | claude · ok · 124.1s |
Notes
Omitted for lack of
A hosted audio source-separation model or API (Demucs/Spleeter class) would enable uploading or selecting an actual recording and using its isolated stems. This implementation synthesizes its four separate sources directly and makes no claim to separate recordings; uploads remain outside the agreed scope.