Evaluation pipeline for a support agent
Production-grade LLM engineering: fast, measurable and cheaper to run.
- Client
- FinTech product team
- Industry
- FinTech
- Services
- LangSmith
- Timeline
- 8 weeks
- Team
- 3 specialists

- +20%Quality
- -40%Model cost
- +20%Output quality score
- 99.9%Uptime
Challenge
A promising prototype that would not scale
Outputs were inconsistent, latency spiked under load and nobody could say whether a prompt change made things better or worse.
Objectives
What we set out to achieve
Make quality measurable and the system reliable enough for production traffic.
- 1Shared evaluation suite
- -30%Latency target
- -30%Cost per request
Strategy
Measure first, then optimise
We built an evaluation suite from real examples, then tuned prompts, models and caching against it.
Solution
An LLM stack built for production
Model routing, caching, structured outputs and automated evals on every change  with dashboards the whole team can read.
Technology
The stack behind it
Chosen for reliability, speed and easy ownership by the client team.
- OpOpenAI
- ClClaude
- LaLangSmith
- PyPython
- FaFastAPI
- ReRedis
- AWAWS
- GrGrafana
Implementation
8 weeks, four phases
Weekly demos kept the client team involved at every step.
- Weeks 1–2Audit & evalsBaseline quality, latency and cost.
- Weeks 3–5Re-architectureRouting, caching, structured outputs.
- Week 6Load testingThroughput and failure modes.
- Weeks 7–8HandoverCI evals, dashboards, runbooks.
Results
Quality you can prove
+20% Quality, and every future change is checked against the same evaluation suite.
