SEA-CLIP Explorer
failedA browser-only image/text matching playground built on SEA-CLIP-Tiny, so anyone can type a word in Vietnamese, Thai, Tagalog, or Bahasa and watch it find the matching image in real time.
Approach
SEA-CLIP-Tiny (fassabilf/sea-clip-tiny, arXiv:2609.30739, CC-BY-4.0, loadable via `open_clip.create_model_and_transforms('hf-hub:fassabilf/sea-clip-tiny')`) is a real open-weight CLIP-style model with a ViT-T-16 image tower and a much smaller text tower, published five days ago with no existing browser port. A one-time Python build step downloads the checkpoint plus ~20 real CC-licensed photos from Wikimedia Commons (food, places, objects recognizable across the five languages), runs the real vision encoder once offline to bake fixed image embeddings into a static JSON file, and separately exports just the text tower to a quantized ONNX file small enough to ship. The shipped page is a single no-bundler index.html: it vendors onnxruntime-web and the model's CLIP-BPE vocab/merges as local static files (no CDN, no Hugging Face calls at runtime) and runs the quantized text encoder client-side in WebAssembly to embed whatever the visitor types. Cosine similarity against the precomputed image embeddings re-ranks the same 20-image grid live, so the interesting claim — does this genuinely align non-English SEA-language text with real photos, not just English — is directly testable with the real model's real weights, mismatches and all, rather than a mocked demo. Skipping the vision tower in-browser is the scope cut that keeps this to about an hour: no image upload, no WebGPU dependency, just a small transformer forward pass over a handful of tokens.
The source post
https://x.com/ashvanth_s1/status/2104533748012237275
Scoring
Pick
| surprise | 4 |
|---|---|
| demonstrability | 5 |
| self_containedness | 5 |
| honesty | 5 |
SEA-CLIP-Tiny is a real, open-weight CLIP model trained specifically on Southeast Asian languages and imagery, released with paper and code on Hugging Face -- it runs entirely client-side, needs no account, no key, and no server, and the interesting claim (does it actually understand Vietnamese/Thai/Tagalog/Bahasa text against images, not just English) is directly testable by typing a word and watching the match happen. Every other candidate tonight either needed something we don't have (a paid image-gen API for the two LoRAs), described infrastructure rather than an artifact (Agent Foundation's runtime, Open-Dots' full computer-use clone), required native mobile apps (MausBot), or needed local filesystem/IDE access that can't be safely exposed as a public deployable page (Lasso, Graphite). SEA-CLIP-Tiny was the only idea that is simultaneously small, self-contained, honestly testable, and not already built.
Review
no scores recorded
Cost
| total | $0.0000 |
|---|
What it looked at
Lasso open source**: a dev tool letting you select any UI element in a running app and have AI edit the real source code; fully inspectable and forkable on GitHub.
Graphite**: open-source (Apache 2.0) tool that converts JVM/Android bytecode into a queryable program graph so AI coding agents can explore codebases with Cypher.
Agent Foundation (a13n)**: self-hosted open-source runtime for agents featuring memory, sandboxes, computer use, and durable execution—ready to run locally.
MausBot**: open-source always-on agents that get their own computers, beautiful iOS/Android apps, and can start tasks locally or in the cloud.
Open-Dots**: fully open-source, self-hostable clone of OpenAI’s Dots with browser/terminal/files computer use and personal app connectors, works with any LLM.
Body Swap LoRA for Qwen Image 2.1**: open-source LoRA that replaces entire people in scenes while preserving pose, framing, and background—weights and examples on Hugging Face.
SEA-CLIP-Tiny**: tiny open-source CLIP model + data tailored for Southeast Asian languages and images, fully released on Hugging Face with paper and code.
Urotsukidoji character LoRA**: free open-source LoRA exploring new versions of a classic anime demon (gym rat, barista, etc.) with full training set on Civitai.
Gate
| pass | project directory exists — /Users/artax/code/builds/2026-09-30/project |
|---|---|
| pass | dependencies install |
| fail | build succeeds — exit 1 |
Notes
Omitted for lack of
A network-enabled build executor with binary downloads from Hugging Face, Wikimedia Commons, PyPI, and the npm registry, plus Python 3.11–3.13, is needed to prepare the real checkpoint, photo collection, quantized encoder, and vendored WASM runtime. No inference API, account, API key, generated illustration, or runtime external service is needed or used.
Notes
gate: FAIL build succeeds — exit 1 (after 2 repairs)