Utkarsh Bali

building Checkpoint (opens in a new tab)

  • As a kid, I drew up quadcopters, sketched Iron Man suits, and built motor-powered cars I insisted on calling “Thrust SSC.”
  • In high school, I started building the apps and games I wished existed.
  • Came to Purdue in 2023 for CS and AI, with a minor in psychology.
  • Started doing AI research there in my second year, and spent a year on a Microsoft collaboration reading a few million social posts about Minecraft.
  • Built WalleX, an AI wallpaper app, around the same time. 3,000+ people in 22 countries used it. The first thing I made that had real users in it.
  • Built a clinical assistant that nurses tested across Indiana hospitals. It cut about 40% of their documentation time, and it runs on the hospital’s own hardware, because patient data should not leave the building.
  • Worked at a San Francisco-based YC startup, QualGent (YC X25), reporting to the CTO, and built agent infrastructure into production for the first time.
  • Co-founded Checkpoint to build agent testing infrastructure. YC told us we were in the top 10% of Summer 2026 applicants, then didn’t interview us. It’s live in private beta.
  • Spent this summer at Recurly, building the thing that automates the company’s Product Development Lifecycle (PDLC). It turns a product requirement into shipped production software. Also, built a Sales outbound automation tool to assist SDRs.
  • Spent the last few months on CLIP-H, a framework that helps generate trustworthy clinical hypotheses. It is going into a NeurIPS workshop submission with Purdue and Harvard Business School faculty.
  • TA’d AI courses twice and ran weekly ML workshops for a few hundred students along the way.
  • Finishing at Purdue in December.

Selected work

  • Recurly Agent Platform

    Takes a product requirement and turns it into merged pull requests. Five agents do the work. A human signs off five times.

    Agent infrastructure · 2026 · ~3x faster from requirement to merge

  • Checkpoint

    CI/CD for AI agents. Adversarial test suites that run before an agent ever meets a user.

    Agent reliability · 2026 · Top 10% of YC Summer 2026 applicants

  • CLIP-H

    A clinical prediction that arrives with its reasoning attached, as named hypotheses a clinician can read and argue with.

    Interpretability research · 2026 · NeurIPS 2026 workshop submission, under review

Where the work went

Get in touch

Also an essay (opens in a new tab) on Medium: What a Quiet Mind Taught Me About God, Family, and Infinity.