01 — Software & Data Platforms · Customer experience

AI ENGINEERING

A test-and-safety harness that let a new agentic product launch on hard evidence.

Up to 5,000

simulated conversations at once, before a real customer saw an agent.
01 — Software & Data Platforms · Customer experience

AI ENGINEERING

A test-and-safety harness that let a new agentic product launch on hard evidence.

Up to 5,000

simulated conversations at once, before a real customer saw an agent.

Industry

Customer-experience technology company, venture backed.

Scale

Enterprise agentic platform for contact centers across voice and digital channels; 1,000+ employees, $1B+ valuation.

Engagement

Implementation: a 6-month embedded build, core build of roughly 12 weeks with 8 to 10 engineers, plus onboarding and standby support.

Industry

Customer-experience technology company, venture backed.

Scale

Enterprise agentic platform for contact centers across voice and digital channels; 1,000+ employees, $1B+ valuation.

Engagement

Implementation: a 6-month embedded build, core build of roughly 12 weeks with 8 to 10 engineers, plus onboarding and standby support.

Industry

Customer-experience technology company, venture backed.

Scale

Enterprise agentic platform for contact centers across voice and digital channels; 1,000+ employees, $1B+ valuation.

Engagement

Implementation: a 6-month embedded build, core build of roughly 12 weeks with 8 to 10 engineers, plus onboarding and standby support.

Industry

Customer-experience technology company, venture backed.

Scale

Enterprise agentic platform for contact centers across voice and digital channels; 1,000+ employees, $1B+ valuation.

Engagement

Implementation: a 6-month embedded build, core build of roughly 12 weeks with 8 to 10 engineers, plus onboarding and standby support.

// THE PROBLEM

Where time went

The launch window was fixed and customers were waiting. Proving an agent ready by hand meant weeks of scripted QA per iteration, with no dependable number to gate a launch on.

// THE BUILD

What we built with them

A synthetic-customer simulator, extensible evaluation engine, guardrail and PII-redaction layer, and a repeatable risk scorecard that the client’s engineers now extend themselves.

// THE PROBLEM

Where time went

The launch window was fixed and customers were waiting. Proving an agent ready by hand meant weeks of scripted QA per iteration, with no dependable number to gate a launch on.

// THE BUILD

What we built with them

A synthetic-customer simulator, extensible evaluation engine, guardrail and PII-redaction layer, and a repeatable risk scorecard that the client’s engineers now extend themselves.

// THE PROBLEM

Where time went

The launch window was fixed and customers were waiting. Proving an agent ready by hand meant weeks of scripted QA per iteration, with no dependable number to gate a launch on.

// THE BUILD

What we built with them

A synthetic-customer simulator, extensible evaluation engine, guardrail and PII-redaction layer, and a repeatable risk scorecard that the client’s engineers now extend themselves.

// PROOF POINTS

Up to 5,000

simulated conversations before a real customer saw an agent

50

evaluations gating quality on day one

2

lighthouse customers onboarded on the fixed timeline

2

lighthouse customers onboarded on the fixed timeline

3-4 weeks

of manual scripted QA replaced per iteration

3-4 weeks

of manual scripted QA replaced per iteration

Up to 5,000

simulated conversations before a real customer saw an agent

Up to 5,000

simulated conversations before a real customer saw an agent

50

evaluations gating quality on day one

50

evaluations gating quality on day one

2

lighthouse customers onboarded on the fixed timeline

3-4 weeks

of manual scripted QA replaced per iteration

STACK

AWS

Claude on Bedrock

LLM-as-judge evals

Synthetic Customers

Launched on schedule; the harness has kept hardening every agent they have shipped since, and their engineers extend it by configuration.
// How we measured

Launch completion against the client’s fixed timeline; scale of simulated testing run pre-launch; QA weeks estimated.

Build with confidence

Talk to us about testing your agents before launch.

Build with confidence

Talk to us about testing your agents before launch.

Build with confidence

Talk to us about testing your agents before launch.

Build with confidence

Talk to us about testing your agents before launch.

© 2026 LevelUp Labs®. All rights reserved.

© 2026 LevelUp Labs®. All rights reserved.