cipris

cipriscipriscipris

cipris

cipriscipriscipris
Abstract digital API technology concept with charts and gears.

Track Record — AI Systems Evaluation

AI Evaluation. Reproducible Evidence.

Selected achievements across code and reasoning evaluation, benchmark design, agent workflows, deterministic verification and evaluation quality control.


Code & Reasoning Evaluation

Mandate
Assess whether AI-generated code and reasoning satisfy the actual specification and constraints.

Delivery
Evaluated correctness, completeness, instruction alignment, code logic and edge cases, identifying failures hidden by otherwise polished outputs.

Outcome
Produced evidence-based evaluations that distinguish substantive correctness failures from presentation or style issues.


AI-Agent Benchmark Production

Mandate
Simplify a fragmented benchmark-production process without weakening controls or auditability.

Delivery
Redesigned the operating workflow from 15 core files to 3, while preserving 10 lifecycle gates and 65 traceable requirements.

Outcome
Reduced duplication and manual reconciliation while retaining controlled human approvals and quality checks.


Professional-Artifact Evaluation & QC

Mandate
Create a consistent evaluation process for AI-generated presentations, documents and spreadsheets.

Delivery
Built a controlled workflow linking task-specific requirements and evidence to verification, findings, scoring, ranking and final QC.

Outcome
Consolidated the evaluation process into 3 reusable workflow components with traceable evidence throughout.


Failure-First Benchmark Design

Mandate
Design document-grounded benchmarks around genuine, reproducible model weaknesses.

Delivery
Tested candidate tasks against the model, retained confirmed failures and excluded ambiguous or already-passing cases, with atomic binary rubrics and explicit expected answers.

Outcome
Converted observed failures — including data-alignment, statistical and classification errors — into reusable benchmark cases.


Quantitative Agent Benchmark

Mandate
Build a realistic quantitative-finance benchmark to evaluate an AI agent’s debugging ability.

Delivery
Built a cross-asset P&L-explain task around four interacting faults spanning dates, signs, scaling and factor mapping, with a reference solution and separate verifier.

Outcome
Passed Idea Check and achieved a local Oracle result of 1.000.


Human-Controlled Evaluation

Mandate
Use automation without removing expert control over material evaluation decisions.

Delivery
Combined deterministic checks, regression controls and explicit human escalation points within the evaluation workflow.

Outcome
Focused expert judgement on genuinely material or ambiguous decisions while keeping routine checks consistent and traceable.


Selected Outcomes

  • Benchmark workflow simplified from 15 core files to 3
  • 65 requirements converted into traceable controls
  • Confirmed model failures turned into reproducible benchmarks
  • Quantitative agent benchmark passed Idea Check and local Oracle validation

Discuss Your Requirements

© 2026 cipris - All Rights Reserved.

  • Home
  • Privacy Policy

This website uses cookies.

We use cookies to analyze website traffic and optimize your website experience. By accepting our use of cookies, your data will be aggregated with all other user data.

DeclineAccept