Shopping Copilot
Built for TikTok TechJam 2026 · Track 4
Offline · Verified
01 Judge-facing evidence tour

Shopping
Copilot

Find the purchased product earlier and rank it higher — across 50,000 catalog items, in at most 10 turns.

Inspect verified results
Deterministic Offline CPU Inspectable traces
Competition Data Contract
What we agreed to solve
Competition Evidence

SCENARIO MIX

Catalog Samples
Text-only metadata — no product images
Step 1 of 6
Scenario Replay
Real public sessions from the official evaluator
Competition Evidence

Route

Hard Constraints

Soft Preferences

Negative Constraints

Ask Attribute

Top-10 Recommendations Initial list
Turn 1
Step 2 of 6
Mechanism Inspection
How the pipeline processes a real request
Competition Evidence
Experiments we did not ship Negative Result
Step 3 of 6
Evaluation Evidence
Official public-set evaluator result · Not hidden-set evidence
Competition Evidence

VERSION COMPARISON

Version HitRate@10 MRR MTTC TechnicalScore

PER-SCENARIO BREAKDOWN (V1.3)

Scenario N HitRate@10 MRR MTTC
Reproduce & artifact provenance Command · report hash · agent commit

          
Step 4 of 6
Beyond the Benchmark — Transparent Ads
Demo-only commercial extension
Demo Only
Demo-only commercial extension.
Never called by the official evaluator. Does not alter organic Top-10 ranking.

PRESET EXPERIMENT

Campaign A

Bid$1.00
Relevance0.82
eCPM

Campaign B

Bid$5.00
Relevance0.12
eCPM
This chapter proves: transparent ad auction · relevance-aware monetization · budget accounting · official/demo path isolation · impact potential.
Current scope: impression auction and budget accounting. No click/conversion path, CTR, or GMV.
Step 5 of 6
Deliverables & Limitations
Competition Evidence
Limitations 5 explicit claim boundaries
  • Public-set gains fix generic error classes but do not substitute for the hidden 800-session validation.
  • Popularity tiebreaker band (5.0) is tuned on public 200; a conservative band=3 (+0.03) is available if hidden set regresses.
  • The optional Qwen layer improves vague-intent classification but is not required for reported scores.
  • No dense/vector recall — proven unnecessary since recall is 100% saturated (200/200).
  • Ad engine is demo-only with simulated inventory and budgets.
What we ship: a deterministic, offline, reproducible scored Agent.
What we demonstrate beyond the score: transparent commercial extensions.
What remains unknown: organizer-private 800-session performance.
Step 6 of 6 — Tour Complete Explore all evidence →