
Selected achievements across code and reasoning evaluation, benchmark design, agent workflows, deterministic verification and evaluation quality control.
Mandate
Assess whether AI-generated code and reasoning satisfy the actual specification and constraints.
Delivery
Evaluated correctness, completeness, instruction alignment, code logic and edge cases, identifying failures hidden by otherwise polished outputs.
Outcome
Produced evidence-based evaluations that distinguish substantive correctness failures from presentation or style issues.
Mandate
Simplify a fragmented benchmark-production process without weakening controls or auditability.
Delivery
Redesigned the operating workflow from 15 core files to 3, while preserving 10 lifecycle gates and 65 traceable requirements.
Outcome
Reduced duplication and manual reconciliation while retaining controlled human approvals and quality checks.
Mandate
Create a consistent evaluation process for AI-generated presentations, documents and spreadsheets.
Delivery
Built a controlled workflow linking task-specific requirements and evidence to verification, findings, scoring, ranking and final QC.
Outcome
Consolidated the evaluation process into 3 reusable workflow components with traceable evidence throughout.
Mandate
Design document-grounded benchmarks around genuine, reproducible model weaknesses.
Delivery
Tested candidate tasks against the model, retained confirmed failures and excluded ambiguous or already-passing cases, with atomic binary rubrics and explicit expected answers.
Outcome
Converted observed failures — including data-alignment, statistical and classification errors — into reusable benchmark cases.
Mandate
Build a realistic quantitative-finance benchmark to evaluate an AI agent’s debugging ability.
Delivery
Built a cross-asset P&L-explain task around four interacting faults spanning dates, signs, scaling and factor mapping, with a reference solution and separate verifier.
Outcome
Passed Idea Check and achieved a local Oracle result of 1.000.
Mandate
Use automation without removing expert control over material evaluation decisions.
Delivery
Combined deterministic checks, regression controls and explicit human escalation points within the evaluation workflow.
Outcome
Focused expert judgement on genuinely material or ambiguous decisions while keeping routine checks consistent and traceable.
© 2026 cipris - All Rights Reserved.