Novarch
Talk to founders

The team

Two founders who built evaluation infrastructure inside Microsoft.

We’ve spent the last two years inside one of the largest enterprise agent rollouts, building the evaluation infrastructure that decides whether Microsoft Copilot products are safe to ship and watching where existing systems break. Novarch turns that discipline into a runtime layer your customers can put in front of their own agents: pinned models, cited rules, structured evidence, and an operator in the loop on borderline cases.

Portrait of Sid Vemuri

Co-founder

Sid Vemuri

CEO

Spent three years as a PM at Microsoft Power BI, where he led the evaluation program for its Copilot and shipped AI features to millions of users. Featured speaker at FabCon 2024, 2025, and 2026.

Peer-reviewed ML researcher with an MS in Machine Learning: first author on a CogSci 2024 paper measuring how well deep-learning models capture human concepts, run across 18 model architectures.

Now
Novarch · Co-founder & CEO
Prior
Microsoft Power BI · PM, led Copilot evaluation
Education
Georgia Tech · M.S. in Machine Learning
Research
CogSci 2024
LinkedIn

Portrait of Sandra Ho

Co-founder

Sandra Ho

CTO

On Microsoft’s Security AI Research team, she builds agent-evaluation infrastructure for the enterprise. The LLM evaluation infrastructure she built there became part of Azure AI Foundry. Production agent-eval infrastructure for an enterprise audience is something very few people in the industry have shipped.

Second author on CTI-REALM, a published agentic security benchmark (March 2026). Microsoft’s EVP of Security cited it publicly when announcing the Project Glasswing collaboration with Anthropic.

Now
Microsoft · Security AI Research (Agent Evals)
Built
LLM eval infrastructure → Azure AI Foundry
Education
Carnegie Mellon University
Research
CTI-REALM
LinkedIn

Advisors

Portrait of Chris Yasko

Chris Yasko

Former SVP & Chief Data Scientist, Equifax

At Equifax he ran global AI/ML research, the data science labs, IP development, and the work on explainable AI. In credit decisioning, a model that cannot say why it decided something is one you cannot put in front of a regulator.

Three decades across high-growth startups and Fortune 100 enterprises: productizing emerging technology, scaling R&D and engineering teams, and shipping under regulatory constraint.

Now
Novarch · Advisor
Prior
Equifax · SVP & Chief Data Scientist
Ran
Global AI/ML research, data science labs, IP
Advises on
Market strategy and commercialization
LinkedIn

Why this team

Eval discipline came first and the product came after.

The load-bearing decisions in Novarch all came from the same place: years spent asking whether an AI system is good enough to ship into a workflow that matters, and seeing evaluation used to react to failures after the fact instead of stopping them before they ship. That’s where one call per action, a pinned model SHA, structured output with a cited rule and signals, a database-rendered audit document, and an operator on borderline cases all come from.

Runtime enforcement is downstream of that question. Novarch is the version of the answer your customers can run on agents you didn’t train.