Evaluation and quality assurance for AI systems in finance and risk.
Independent research consultancy. We design evaluation tasks, grading rubrics and review processes for language models operating in quantitative finance, counterparty risk and market infrastructure.
CONTACT ↓01What we do
01.1Evaluation design
Task sets, benchmarks and scenario libraries that test whether a model reasons correctly about derivatives, collateral, margin, credit exposure and regulatory logic — not whether it sounds fluent.
01.2Rubrics and grading
Scoring frameworks that a non-specialist can apply consistently, with domain-expert ground truth behind every rule. Built for RLHF pipelines, red-teaming and model comparison.
01.3Domain QA and review
Expert review of model outputs, synthetic data and training material in quantitative finance and risk. We catch the errors that look right to a generalist.
02How we work
02.1Domain first. Every task is written by someone who has run the underlying process at a bank or central counterparty, not adapted from a textbook.
02.2Reproducible. Rubrics are versioned, testable and documented. A grader six months from now gets the same answer.
02.3Confidential by default. Client work is never referenced, named or reused. Engagements are covered by NDA as standard.
02.4Remote, asynchronous. We work with research and data teams across US, European and Asian time zones.
03Why "second line"
In risk management, the first line builds and runs the business. The second line monitors and challenges it. We are the second line for AI systems: independent evaluation, challenge and assurance, applied by people who have done the first-line job.
04Background
The practice draws on more than a decade in front-office and supervisory risk roles — counterparty credit risk and central-counterparty quantitative research at a global investment bank, and CCP risk supervision at a G7 central bank. Areas of depth include margin models, stress testing, exposure simulation, model validation and clearing-house risk frameworks.
The practice focuses on evaluation and quality assurance for frontier AI systems in these domains.