← Zalman Friedman

Agentic AI, MCP, and Evals in Legal Tech

Propela (contract), Apr–Jul 2026

Problem
An AI legal-tech platform needed to move from assistive AI to an agentic system that could act across an expert workflow — case-data extraction, drafting, editing, and exporting — reliably and measurably.
My role
Product Strategist (contract). Drove the agentic platform, a custom MCP server, and the AI evaluation framework.
Outcome
We accelerated a manual client workflow by ~90%, and reduced another team's staffing needs by ~60%. The MCP server and eval framework were built and in place at engagement's end.

Situation

Letting an AI system act on legal workflows raises the two questions every agentic product has to answer: what is the agent allowed to touch, and how do you know it's good enough? Most of my work was building the product infrastructure that answers them.

The decisions

The agentic pipeline. The platform autonomously handled case-data extraction, drafting, editing, and exporting, an expert workflow that had taken a trained reviewer ~45 minutes per case. The product question at every step was what the agent could complete alone vs which required human review before anything irreversible (like a filing) happened.

The custom MCP server. I drove a Model Context Protocol server connecting external AI agents to clients' case-management systems through Propela's authentication layer. The PM work was defining the tool surface: which operations agents get, at what granularity, behind which trust boundaries. A too-broad surface is a liability; a too-narrow one makes the agent useless.

The eval framework. Non-deterministic output needs a defined quality bar you're shipping against. I drove an evaluation framework measuring the accuracy of the agentic system against a non-agentic golden sample baseline, defining the success metrics and what "good enough to act autonomously" means. (Results were still being collected when the engagement ended.)

Delivery cleanup. Alongside the AI work, I imposed definitions of done, scoped estimates, and prioritization on roughly ten open workstreams across a 10-person engineering team, cutting active workstreams by 60%.

What it shows

End-to-end ownership of the three hardest AI-platform surfaces — agent design, protocol/tool-surface scoping, and evaluation, plus a clear point of view on where agents are reliable and where a human must stay in the loop. I also build with agents daily on my own projects, so the judgment comes from real world practice.