Skip to content

Work

Everything I’ve built worth writing up.

7 projects, from production agent infrastructure to research prototypes. Each one covers the problem, the architecture, and what it actually proved.

Agent reliability2026

Top 10%

of YC Summer 2026 applicants

Agent reliability · 2026

Checkpoint

CI/CD for AI agents. Adversarial test suites that run before an agent ever meets a user.

Read the case study

Interpretability research2026

0.844

AUROC against a synthetic oracle

Interpretability research · 2026

CLIP-H

Clinical hypothesis verification on MIMIC-IV using sparse autoencoders and an LLM ensemble.

Read the case study

Agent infrastructure2025

<1%

task failure rate at scale

Agent infrastructure · 2025

App Crawler

A GKE-distributed crawler that indexes Android apps into an org-wide knowledge base for downstream agents.

Read the case study

More work

  • QualGent AI Assistant

    An enterprise QA copilot routing across 45+ tools and sub-agents from a single orchestrator.

    QualGent (YC X25) · 2025

  • LLM Multi-Agent QA System

    Four agents (planner, executor, verifier, supervisor) driving Android UI flows to deterministic completion.

  • Clinical AI Assistant

    Self-hosted LLaMA and on-device speech, tested across Indiana hospitals under HIPAA constraints.

    Purdue University · 2024 to 2025

  • WalleX

    An AI wallpaper app with nine image models. My first product with real users in it.

Contact

Building something in this world? Let’s talk.

I’m always up for a conversation about agents, developer tools, or a product you think should exist. The fastest way to reach me is email.

baliutkarsh2@gmail.com