LLM evaluation and benchmarking

Evaluation that says which model is strong where

A single leaderboard number tells you almost nothing about whether a model can do your work. We run the same operational task across a model set under matched conditions and score every run on more than eight dimensions, so the result is a per dimension comparison rather than one figure.

Active in the micro1 Company Data Partnerships Program.

How we evaluate

Comparability is the whole game. If the prompt, the context and the harness are not held constant, the differences you measure are differences in your setup.

Matched conditions, every model

The same task, the same context, the same harness, the same acceptance criteria. When something separates two models, that separation is attributable to the models. This is the part most internal evaluations get wrong, because the prompt drifts as the team learns what works.

An expert baseline to score against

Every task carries a completion by a practitioner who does that work for a living. Without a human ceiling, a rubric measures models against each other and never against good. With one, you can say whether the gap that remains actually matters.

Ranked per dimension, not averaged

More than eight dimensions, agreed with you before work starts, applied identically across the set. Averaging them away hides the finding: models rarely fail uniformly, they fail in a shape. We report the shape.

What never leaves our side of the table

This work contributes operational expertise, not data about the people we work for. These are operating rules, applied before a workflow is written rather than checked afterwards.

Scrubbed or fictionalised from the start

Where a real workflow touches anything sensitive, it is rebuilt as a scrubbed or fictionalised version before any work begins. Not redacted afterwards.

Never client or customer material

No client, customer or employer confidential information. No proprietary material. Nothing covered by an NDA. This is the same commitment that governs our software delivery work.

No personal data, no credentials

No private or personal information about anyone, and no passwords, API keys or credentials in any prompt, file or environment.

Only tools we can properly grant

We work only in tools where the operator personally holds access and can safely extend that access to an AI system. Where a real process would run through a company system, we substitute a personal or fictionalised equivalent.

Never on unauthorised devices

Agent work does not run on any device we are not authorised to run it on.

The functions we can cover

Breadth from running these functions as a working software business, not from staffing up for a programme. We are also among the top one percent of official FlutterFlow partners globally, which is what betting early on a shift looks like when it pays off.

Software engineeringQA and testProduct and project deliveryCustomer supportSales and CRM operationsInternal operations

How an engagement runs

1

Scope

We agree the functions, the difficulty bar and the rubric dimensions with you before anything is built.

2

Build

Workflows are authored by the people who do the work, then scrubbed or fictionalised and checked against the data rules above.

3

Run

Each workflow is executed across the model set under matched conditions, with the expert baseline captured alongside.

4

Score and compare

Every run is graded on the agreed dimensions, then models are ranked against each other so the per dimension picture is legible.

5

Deliver

You receive the workflows, the runs, the scores and the comparison, in the structure your pipeline expects.

Questions partners ask

How is this different from a public benchmark?

Public benchmarks are saturated, contaminated by training data, and built from tasks that look nothing like operational work. Ours are authored from functions we run as a working business, they have never been published, and they are hard enough that current frontier models struggle with them.

Can you evaluate against our own rubric?

Yes, and we would rather. If you already have dimensions your pipeline expects, we work to those instead of ours. Where you do not, we agree the rubric with you before anything is built so the results are usable on arrival.

Which models do you cover?

The set is agreed per engagement, and we run whatever is in it under identical conditions. We do not publish which models we have evaluated or the results, and would extend the same discretion to your programme.

Who does the work?

The same practitioners who do the job for clients. The person who defines a QA workflow is a QA practitioner, and the person grading a software engineering run is an engineer. That is the point: the expert baseline has to be genuinely expert or the comparison is not worth much.

Do you ever use real client data?

No. Client confidentiality and intellectual property are non negotiable, and they do not bend for this programme. Where a workflow is drawn from real operational experience, it is rebuilt as a scrubbed or fictionalised version before the work starts, not cleaned up afterwards.

How do we start?

A short call to scope one batch. We would rather prove the work on a small, well defined set than talk about volume before you have seen the quality.

Related

Bring us the rubric, we will bring the tasks

One well defined batch is a better conversation than a capability deck. Tell us the functions you care about and the dimensions you score on.

Start a conversation