AI training data

AI training data from work that actually happens

Empiric Infotech builds hard operational workflows for frontier AI programmes, runs them across multiple models under matched conditions, and scores every run against a rubric agreed with you, dimension by dimension, with written feedback behind each score. The tasks come from functions we run as a working software business, so the difficulty is real rather than invented.

A service we provide to AI labs, data partnership programmes and companies training or fine tuning their own models. Active in the micro1 Company Data Partnerships Program.

What we contribute

Programmes do not need more easy tasks. They need work that separates one model from another, and a credible human baseline to separate it against.

Workflows built to break the model

We take real operational work and turn it into tasks that are genuinely hard for a frontier model: long horizons, ambiguous inputs, branching decisions, and the judgement calls that only show up in production. A task earns its place only if it is hard for the right reason, which is that the work itself is difficult rather than that the brief was badly written.

The same workflow, across several models

One task run once against one model tells you almost nothing. We run the same workflow across multiple frontier models under the same conditions, so the differences that surface are differences in the models rather than differences in the prompt.

Scored per dimension, with the reasoning

Each run is graded against a rubric agreed with you before work starts, dimension by dimension, and every score carries written notes on what the model handled and where it went wrong. The output is not a single number. It is a map of which model is strong where, and why.

What never leaves our side of the table

This work contributes operational expertise, not data about the people we work for. These are operating rules, applied before a workflow is written rather than checked afterwards.

Scrubbed or fictionalised from the start

Where a real workflow touches anything sensitive, it is rebuilt as a scrubbed or fictionalised version before any work begins. Not redacted afterwards.

Never client or customer material

No client, customer or employer confidential information. No proprietary material. Nothing covered by an NDA. This is the same commitment that governs our software delivery work.

No personal data, no credentials

No private or personal information about anyone, and no passwords, API keys or credentials in any prompt, file or environment.

Only tools we can properly grant

We work only in tools where the operator personally holds access and can safely extend that access to an AI system. Where a real process would run through a company system, we substitute a personal or fictionalised equivalent.

Never on unauthorised devices

Agent work does not run on any device we are not authorised to run it on.

The functions we can cover

Breadth from running these functions as a working software business, not from staffing up for a programme. We are also among the top one percent of official FlutterFlow partners globally, which is what betting early on a shift looks like when it pays off.

Software engineeringQA and testProduct and project deliveryCustomer supportSales and CRM operationsInternal operations

How an engagement runs

1

Scope

We agree the functions, the difficulty bar and the rubric dimensions with you before anything is built.

2

Build

Workflows are authored by the people who do the work, then scrubbed or fictionalised and checked against the data rules above.

3

Run

Each workflow is executed across the model set under matched conditions, so any difference that surfaces is attributable to the model rather than the setup.

4

Score and compare

Every run is graded on the agreed dimensions, then models are ranked against each other so the per dimension picture is legible.

5

Deliver

You receive the workflows, the runs, the scores and the comparison, in the structure your pipeline expects.

Questions partners ask

Do you ever use real client data?

No. Client confidentiality and intellectual property are non negotiable, and they do not bend for this programme. Where a workflow is drawn from real operational experience, it is rebuilt as a scrubbed or fictionalised version before the work starts, not cleaned up afterwards.

Is this data annotation or labeling?

No, and the distinction matters. Annotation work is priced by the hour against very large offshore labeling operations, and volume is the product. What we produce is expert authored operational tasks, comparative scoring across models, and written feedback from a practitioner on what each model handled and where it failed. The unit is a workflow and a judgement, not a label.

What makes a workflow worth evaluating?

Difficulty that is real rather than artificial. The task has to reflect how the work is genuinely done, it has to be hard enough that current frontier models struggle with it, and an expert in that field has to be able to complete it with precision. A task that everything passes tells a programme nothing.

Who does the work?

The same practitioners who do the job for clients. The person who writes a QA workflow is a QA practitioner, and the person scoring a software engineering run is an engineer. That is the point: whoever judges the run has to know what good looks like in that function, or the score and the feedback are not worth much.

Which functions can you cover?

Software engineering, QA and test, product and project delivery, customer support, sales and CRM operations, and internal operations. The breadth comes from years of running these functions as a software delivery business, not from staffing up for a programme.

Can I join as an AI trainer, or can my company supply this to you?

Neither, and it is worth being direct about it. This page describes a service we sell to companies, not a programme you can join and not work we subcontract out. We are not recruiting AI trainers, evaluators or annotators, and we do not buy this capacity from other agencies, because the whole point is that the people writing and scoring the workflows are our own practitioners doing that job for clients. If you are looking for a role at Empiric, our open positions are at /career.

How do we start?

A short call to scope one batch. We would rather prove the work on a small, well defined set than talk about volume before you have seen the quality.

Related

The businesses that built the workflows should help shape the models

If you are sourcing hard operational tasks or comparative model evaluation, we would rather show you one batch than talk about volume. Tell us the functions and the rubric you work to.

Start a conversation

SCOPE A BATCH

Tell us the functions you care about and the rubric you score on, and we will come back with what a first batch would look like.

Phone
0 / 1000
Attach a filePDF, DOC, or image. Maximum 10 MB.

We respond within one business day. Your details stay confidential.