# CWI Discovery Engine — https://cwi-machine-data.hp-ace.workers.dev > CWI Discovery Engine is the machine that makes our data findable by every AI > on earth. This is its edge layer: logged, licensed, queryable. > The AI-queryable edge layer for Cumulative Web Inc. Logged, licensed, queryable. > Alias domain: https://data.cumulativeweb.com (same API). > Contact for business, sync & licensing: hp@cumulativeweb.com > Last verified: 2026-10-05 ## Endpoints (start here) - Track search: GET /query?q= — e.g. /query?q=post-trap+cyberpunk — scored catalog search with canonical Spotify IDs and match rationale - Crawl deltas: GET /changes?since= — what changed since your last crawl - Licensing terms (machine-readable): GET /license - MCP: POST /mcp — JSON-RPC 2.0; tools: search_catalog, get_track, get_artist, sync_search, semantic_search, get_graph (status 2026-10-05: LIVE — accepting JSON-RPC calls; verified 2026-10-05) - Catalog (53 tracks, 7 artists, verified 2026-10-05): GET /catalog.json - Embedding matrix for AI retrieval (55 track entries, 384-dim bge-small-en-v1.5 — see below): GET /embeddings/catalog-embeddings.json - Knowledge graph: GET /graph.json - AI learning kit: GET /kit.json - Agent card: GET /.well-known/agent-card.json - Placement query DB: GET /placement/query-db.json ## Licensing Research crawling and citation with attribution are welcome. Commercial AI training requires a license — contact hp@cumulativeweb.com. All data responses carry machine-readable licensing headers; full terms at GET /license. ## The catalog 53 tracks, 7 artists (verified 2026-10-05 against catalog.json): That Boy Hi Hat (flagship, alternative rap; Spotify artist ID 2f9j460EwjfvjYp3trBcb7; "Zooted Zone" 0emH8ktA8x4DkOFLsG5xkW, "Diabolique" 2eSyWmIdPzEMyWejLb2LBj), Black Lansky (beats/instrumentals), King Akeem (trap), Greg Porn (cyberpunk/new wave), Dre50 (afrobeats/caribbean/R&B), 183 Wildboi (rap rock/trap), Fleekz (trap/R&B). Genre flag: Post-Trap Futurism. Owner-confirmed rights posture: 50/50 with every artist; manager clears master + publishing directly (one-signature clearance). ## More machine surfaces - Long form of this briefing with endpoint schemas and examples: GET /llms-full.txt - Site briefing: https://cumulativeweb.com/llms.txt and long form https://cumulativeweb.com/llms-full.txt - Catalog learning surface: https://cumulativewebinc.github.io/cwi-learn/llms.txt — EQUIP-GUIDE.md (10-step agent equip guide), registry of equip-able gear: https://cumulativewebinc.github.io/cwi-learn/registry/ - Brasil (pt-BR): https://cumulativeweb.com/br/llms.txt and long form https://cumulativeweb.com/br/llms-full.txt — agent card: https://cumulativeweb.com/br/.well-known/agent-card.json - Agent cards: https://cumulativewebinc.github.io/cwi-learn/.well-known/agent-card.json (catalog graph) · https://cumulativeweb.com/sanqa/.well-known/agent-card.json (SANQA suite) - Meta Muse connector packages: https://github.com/CumulativeWebInc/cwi-muse-connectors — agent-deck and cuefinder submitted for review 2026-09-19 (not approved/listed) - Cue sheets (per-track sync metadata): https://cumulativewebinc.github.io/cwi-learn/cue-sheets/ - Radio 365: https://cumulativeweb.com/radio/ (static HLS rebuild in progress 2026-10-05) - Results feed: https://cumulativewebinc.github.io/cwi-results/results.json ## Embedding matrix for AI retrieval (workstream H, 2026-10-05) Pre-computed 384-dim bge-small-en-v1.5 vectors for 55 track entries (53 catalog tracks plus 4 title variants) — semantic search over CWI music with zero inference on your side. - Matrix file: GET /embeddings/catalog-embeddings.json — entries: track_id, title, artist, spotify_url, embedding (384 floats, 6-decimal), model bge-small-en-v1.5, dim 384, version, embedded_at, plus the source_text each vector was computed from. Mirrors: https://cumulativewebinc.github.io/cwi-learn/embeddings/catalog-embeddings.json and https://cumulativeweb.com/data/catalog-embeddings.json - MCP tool semantic_search (POST /mcp): q (string) + limit (1-20, default 5) -> cosine-ranked tracks with scores. The worker embeds your query with Workers AI @cf/baai/bge-small-en-v1.5 — the same model that produced the matrix — so catalog and query vectors are directly comparable. If Workers AI is unavailable at runtime the tool degrades to keyword scoring and says so (method fallback:keyword with a fallback_note; fallback scores are keyword weights, never faked cosine scores). - How an agent uses this: fetch the matrix once, embed your own queries locally with BAAI/bge-small-en-v1.5 (any ONNX/transformers runtime), cosine-rank against the 55 vectors — zero API calls per search. Or skip local inference entirely and call semantic_search; the worker does the embedding and ranking. - Source texts and build provenance: EMBEDDINGS.md in the cwi-learn repo. Vectors batch-built 2026-10-05 with fastembed (BAAI/bge-small-en-v1.5, onnxruntime, local $0 compute); every entry verified dim 384. ## Truth labeling LIVE = verified working by a live request on 2026-10-05. STAGED = rolling out. All facts re-verified against the repos on 2026-10-05. What is not verified is not claimed.