Still in Love
Palo Alto, CA
AI Product Builder
Co-founded an AI product studio, building with ex-Googlers and Stanford OVAL. Runs agentic development across 12 models tiered by capability and risk; keeps development sandboxed from production and gates releases on shadow traffic, daily evals, and human-approved rollouts.
YeahRotten Tomatoes for YouTube reviews (running shoes as a proof of concept)
- Set the quality bar with human judgment: multi-model consensus and adversarial review handle volume; verdicts are calibrated against 20 years in the category and field testing with a 100-member run club.
- Rejected an 86× cost cut to preserve trust: deterministic-rubric evals showed 45% fewer verified claims per video (11.9 → 6.6) at $0.005 vs. $0.43 per inference run; kept the higher-cost path and set a risk-tiered model policy with daily yield evals.
- Defined the product around a 30-second decision: distilled 4,516 review videos into 26,056 source-linked claims across 3,442 SKUs, so a buyer gets a buy, skip, or consider verdict at a glance.
AtomiciOS voice journal that compounds intent
- Day-7 retention +18%: on-device intent routing cut capture friction and time-to-first-token by 90 ms, so users speak instead of picking a skill or model; shipped in 163 TestFlight builds.