AI data partnerships

We build the workflows that break frontier models

Empiric Infotech contributes hard, real operational workflows and multi model comparative evaluation to frontier AI data partners. We find where models fail at work a human expert does with precision, then score every model on more than eight dimensions so you know which one is strong where.

Active in the micro1 Company Data Partnerships Program.

What we contribute

Labs do not need more easy tasks. They need work that separates one model from another, and a credible human baseline to separate it against.

Workflows built to break the model

We take real operational work and turn it into tasks that are genuinely hard for a frontier model: long horizons, ambiguous inputs, branching decisions, and the judgement calls that only show up in production. The bar is deliberate. A task is only useful if an expert in that field completes it with precision and the model does not.

The same workflow, across several models

One task run once against one model tells you almost nothing. We run the same workflow across multiple frontier models under the same conditions, so the differences that surface are differences in the models rather than differences in the prompt.

Scored on more than eight dimensions

Each run is graded against a rubric covering more than eight separate dimensions of performance, then models are ranked against each other dimension by dimension. The output is not a single score. It is a map of which model is strong where, which is the question a lab actually needs answered.

What never leaves our side of the table

This programme contributes operational expertise, not data about the people we work for. These are operating rules, applied before a workflow is written rather than checked afterwards.

Scrubbed or fictionalised from the start

Where a real workflow touches anything sensitive, it is rebuilt as a scrubbed or fictionalised version before any work begins. Not redacted afterwards.

Never client or customer material

No client, customer or employer confidential information. No proprietary material. Nothing covered by an NDA. This is the same commitment that governs our software delivery work.

No personal data, no credentials

No private or personal information about anyone, and no passwords, API keys or credentials in any prompt, file or environment.

Only tools we can properly grant

We work only in tools where the operator personally holds access and can safely extend that access to an AI system. Where a real process would run through a company system, we substitute a personal or fictionalised equivalent.

Never on unauthorised devices

Agent work does not run on any device we are not authorised to run it on.

The functions we can cover

Breadth from running these functions as a working software business, not from staffing up for a programme. We are also among the top one percent of official FlutterFlow partners globally, which is what betting early on a shift looks like when it pays off.

Software engineeringQA and testProduct and project deliveryCustomer supportSales and CRM operationsInternal operations

How an engagement runs

1

Scope

We agree the functions, the difficulty bar and the rubric dimensions with you before anything is built.

2

Build

Workflows are authored by the people who do the work, then scrubbed or fictionalised and checked against the data rules above.

3

Run

Each workflow is executed across the model set under matched conditions, with the expert baseline captured alongside.

4

Score and compare

Every run is graded on the agreed dimensions, then models are ranked against each other so the per dimension picture is legible.

5

Deliver

You receive the workflows, the runs, the scores and the comparison, in the structure your pipeline expects.

Questions partners ask

Do you ever use real client data?

No. Client confidentiality and intellectual property are non negotiable, and they do not bend for this programme. Where a workflow is drawn from real operational experience, it is rebuilt as a scrubbed or fictionalised version before the work starts, not cleaned up afterwards.

What makes a workflow worth evaluating?

Difficulty that is real rather than artificial. The task has to reflect how the work is genuinely done, it has to be hard enough that current frontier models struggle with it, and an expert in that field has to be able to complete it with precision. A task that everything passes tells a lab nothing.

What do the evaluation dimensions cover?

More than eight, agreed with the partner before work starts, and applied consistently across every model in the set so results stay comparable. We are happy to work to your existing rubric rather than ours.

Who does the evaluation?

The same practitioners who do the work. The person who defines a QA workflow is a QA practitioner, and the person grading a software engineering run is an engineer. That is the point: the expert baseline has to be genuinely expert or the comparison is not worth much.

Which functions can you cover?

Software engineering, QA and test, product and project delivery, customer support, sales and CRM operations, and internal operations. The breadth comes from years of running these functions as a software delivery business, not from staffing up for a programme.

How do we start?

A short call to scope one batch. We would rather prove the work on a small, well defined set than talk about volume before you have seen the quality.

The businesses that built the workflows should help shape the models

If you are sourcing hard operational tasks or comparative model evaluation, we would rather show you one batch than talk about volume. Tell us the functions and the rubric you work to.

Start a conversation