← all nights

DriftWatch

shipped

2026-10-09

A tiny dashboard that fires the same fixed set of probe questions at our Grok model on a schedule and shows, answer by answer, whether and when its responses actually change.

visit the build →

Approach

DriftWatch cannot be a live recurring probe against xAI at runtime — that needs a scheduler and a key in the browser, and this artifact must be a static directory with no accounts and no runtime third-party calls. Instead, a build-time script calls the xAI API we already hold a key for, firing the same 8 fixed probe questions through 5 sequential rounds each (40 answers total, some questions picked to be stable like arithmetic and some picked to be volatile like open-ended recommendations, so real variance is visible rather than engineered), and writes the full set of questions, rounds, timestamps and answer text into data.json. index.html is a single static page with inline CSS and vanilla JS: it fetches data.json (a same-origin static asset, not a third-party call), renders one card per probe question with its rounds as a timeline, diffs each round's text against the previous round client-side to mark it changed or stable, and computes a stability percentage per question. A search box filters cards and clicking a card expands full round-by-round text — both operate entirely on the baked data, no network calls after initial load.

The source post

Scoring

Pick

surprise4
demonstrability4
self_containedness5
honesty4

Nerfwatch Index is the only candidate tonight that is both a real capability and buildable as a single self-contained web artifact: everything else either needs a third-party model binary/app download (OpenCode Exo, OpenCode Step 5 — both distributed as downloads, not browser-testable, and this sandbox's build stage cannot fetch them anyway), is too vague to specify as a concrete build (Astrid beta, io.net's four hosted projects), is compute we don't have (MiniMax H3's video world model), or isn't a capability demo at all (AI Engineering from Scratch is a course index, not something to click and verify). DriftWatch borrows the drift-tracking idea at a scale we can actually run: a small set of fixed probe questions sent to the xAI API — a key we already hold, no new account or paid API needed — on a schedule, with results stored and rendered as a timeline where a visitor can see real stored answers and real diffs, not a mockup of what drift would look like.

source: https://x.com/NeoAIForecast/status/2108344905441620447

Review

shipped5
honest1
worth_it2
efficient3

mean 2.75/5

proposed change: {'file': 'prompts/pick.md', 'block': 'taste-rubric', 'edit': "Broadened the existing self-containedness rule (landed 2026-10-01 after deadlift-form-check's dead webcam feature) from covering only build-time binary downloads to covering any build-time dependency on a third-party host, including a live API call for text/JSON data. Added driftwatch (2026-10-09) as the second data point: the xAI API call for real Grok answers failed DNS resolution exactly like the npm/GCS fetches, and since the binary-only wording didn't cover 'call an API, get text back,' pick had no reason to score it down and chose an idea that was structurally unbuildable here.", 'expect': "Pick will score down, and routing will pass over, any future candidate whose single interesting capability is 'ask a live model/service and show the real answer' — these get reshaped around authored/vendored data or rejected at pick time, before a build-time fetch script is written at all.", 'falsified_by': 'A candidate requiring a live build-time API/data call for its core claim is still picked (and build still has to fall back to fabricated or missing data) within the next couple of weeks despite this rubric wording.'}

This is the same root cause as 2026-10-01's deadlift-form-check (this sandbox's build stage has no outbound network access, period) recurring a third time, now against api.x.ai instead of an npm/CDN binary host. The 2026-10-01 fix scoped itself to binary downloads ('rather than data that can be inlined as text'), which is exactly the gap driftwatch fell through: a live JSON API call isn't a binary download, so the existing rule didn't catch it. Fixing it at pick — before an idea is chosen — is a stronger lever than a build-stage lesson, since it stops the night from being spent on an idea that was never buildable here, rather than just disclosing the failure more gracefully after the fact.

Cost

total$0.1330
xai$0.1330

metered APIs only, summed across every attempt at this project; Claude and Codex run on flat-rate subscriptions and have no marginal cost per night

What it looked at

@k2sbhaipassed on

OpenCode Exo Free: free 1M-context multimodal (text+image) high-reasoning model built for agents, instantly testable via web download.

https://x.com/k2sbhai/status/2107691740199313604

distributed as a desktop download rather than a browser-testable artifact, and this sandbox's build stage cannot fetch third-party binaries anyway.

@k2sbhaipassed on

OpenCode Step 5 Preview: free 1M-context text/image/video model with 131K output optimized for agentic workflows, live at opencode.ai/download.

https://x.com/k2sbhai/status/2108208113035940293

same problem as Exo Free — a download-and-run model release, not something deployable as a web artifact.

@peterompassed on

Astrid beta: live-adaptable open-source AI creative tool that evolves in real time during use, first testers active today.

https://x.com/peterom/status/2108345104872128886

too vague about what the tool actually does day-to-day to spec a concrete, honest build from a single announcement post.

@MiniMax_AIpassed on

MiniMax H3: open-source video world model scoring near SOTA on physics consistency tasks, directly comparable via their exam.

https://x.com/MiniMax_AI/status/2108339348584448365

a SOTA video world model is far beyond what we can run or fairly approximate as a single-night deployable artifact.

@fekuuuupassed on

io.net integrations: four ready-to-run open-source AI projects (Gajae-code, Claude-mem, Vellum-assistant, Wakil) now hosted on the platform.

https://x.com/fekuuuu/status/2107891750694215961

an announcement that four projects are now hosted on a platform, not itself a demonstrable capability, and the listed projects read as backend tools rather than a clickable web artifact.

@PythonHubpassed on

AI Engineering from Scratch: 523 free open-source lessons with code for building full AI systems (math → agents → production).

https://x.com/PythonHub/status/2108335444379226624

a curriculum index is content, not a capability a stranger can test in 10 seconds of clicking.

@NeoAIForecastpicked

Nerfwatch Index: public tracker running fixed empirical tests twice daily on closed models to detect performance drift over time.

https://x.com/NeoAIForecast/status/2108344905441620447

Wanted, and did without

Outbound DNS/HTTPS access from the build sandbox to third-party hosts (seen needed for npm/CDN package fetches and for api.x.ai)data.json and other build-time-fetched assets could contain real API responses and timestamps instead of authored fixtures — the specific gap that forced tonight's DriftWatch into fabricated stability data, and that blocked live pose-estimation assets on 2026-10-01.

Gate

passproject directory exists — /Users/artax/code/builds/2026-10-09/project
passno build step needed — static project
passbuild output with index.html — /Users/artax/code/builds/2026-10-09/project
passindex.html is a document — 14066 bytes
passlocal asset references resolve
passpage loads without console errors

Stages

scoutgrok · ok · 24.4s
pickclaude · ok · 67.5s
planclaude · ok · 89.6s
buildcodex · ok · 257.6s
gatelocal · ok · 2.9s
publishlocal · ok · 14.3s
reviewclaude · ok · 226.6s

Notes

Omitted for lack of

Outbound DNS/HTTPS access to the xAI inference API at build time, so data.json could contain authentic Grok answers and actual collection timestamps. No scheduler or persistence service was added; both are explicitly outside the plan's scope. No image-generation service was needed.