Back to blog

The AI attribution maturity model

Teams can progress from ad hoc AI disclosure to governed work records without collecting every prompt or deploying every integration at once.

WhoWorked team8 min read
Layered translucent forms moving across a dark field.

Teams often approach AI attribution as an instrumentation project.

They list every tool in use, investigate provider APIs, debate browser extensions, and ask whether they need to capture prompts. The project expands before the team has agreed on a more basic question: what decision should the attribution record support?

A consultancy trying to answer a client disclosure request has different evidence needs from an engineering organization measuring model cost. A regulated team proving human review needs different controls from an agency learning which workstreams are changing fastest.

A useful approach starts with shared terms, then improves the reliability of evidence as the team learns what its decisions require.

Maturity model

Better evidence, one level at a time

  1. Level 1

    Ad hoc

    Memory and scattered notes

  2. Level 2

    Structured

    Shared fields and definitions

  3. Level 3

    Connected

    Tool evidence supports declarations

  4. Level 4

    Reconciled

    Events link to real work

  5. Level 5

    Governed

    Rules match client and risk

Mature attribution collects enough evidence for the decision, not the maximum data available.

Maturity has five dimensions

Before looking at levels, it helps to define what is becoming more mature.

Capture

How does the team learn that AI was materially involved? Through memory, a form, an integration, a provider feed, an agent event, or another signal?

Linkage

Can the AI contribution be connected to a person, project, workstream, and delivered outcome, or does it remain an isolated usage event?

Review

Can a responsible person confirm, correct, dismiss, or verify the record?

Governance

Are approved tools, disclosure rules, retention, privacy, and evidence requirements defined for the work?

Reporting

Can the organization use the record to answer a real question about delivery, billing, cost, client trust, capacity, or policy compliance?

A team can be advanced in one dimension and weak in another. Rich provider data with no project linkage is technically detailed but operationally immature. A simple manual disclosure tied to a client deliverable may be more useful.

Level 1: Ad hoc disclosure

At the first level, AI use is discussed but not recorded consistently.

One person adds a note to a document. Another mentions the model in a project channel. A third assumes that AI is just another tool and says nothing. When a client asks what happened, the team reconstructs the answer from memory.

Typical signals

  • No shared definition of material AI use
  • Disclosures vary by person and project
  • Tool usage is visible only in separate vendor dashboards
  • Human review is assumed rather than recorded
  • Client questions trigger a manual investigation

Main risk

The team cannot distinguish “we believe this was reviewed” from “we have a record that it was reviewed.” The absence of structure creates both under-disclosure and overreaction.

Move to the next level

Agree on a minimal vocabulary: involvement level, contribution type, accountable owner, review status, project, workstream, and outcome.

Level 2: Structured self-reporting

At level two, the team records material AI involvement using a consistent form or template.

This is sometimes dismissed as manual or imperfect. It is also the fastest way to establish the operating model before investing in integrations.

Typical capabilities

  • Shared definitions and examples
  • A small set of involvement levels and contribution types
  • Named ownership and review
  • Records attached to projects and workstreams
  • A client disclosure generated from the internal record

Main risk

Self-reporting is easy to forget and hard to enforce. People may interpret categories differently. The record says what someone remembers or declares, not necessarily everything that occurred.

Move to the next level

Identify the highest-value evidence gaps. Connect one or two important signals instead of trying to instrument every AI tool.

Level 3: Connected evidence

At level three, structured declarations are supported by signals from the tools where work occurs.

Evidence can arrive through several paths:

  • A host application posts an event when a model is used.
  • An AI agent reports a completed activity through an API or MCP tool.
  • A provider usage API supplies model, token, cost, and timestamp data.
  • An IDE or repository integration records relevant coding activity.
  • A privacy-preserving extension records that an approved AI tool was active during a work period.

The goal is not to crown one source as universally best. Each source has a different balance of fidelity, coverage, deployment effort, and trust.

Typical capabilities

  • Tool and model evidence supplements self-reporting
  • Usage cost can be connected to a team or project
  • Missing declarations can be surfaced for review
  • Evidence provenance is recorded
  • The team collects less content because metadata is sufficient for many decisions

Main risk

Connected data can create false confidence. A provider log proves that a model call occurred. It does not prove what business outcome resulted or whether a person reviewed the output. An agent that self-reports may produce rich context but remains a cooperative source.

Move to the next level

Reconcile events with the work record. Make uncertainty and missing linkage visible instead of silently discarding awkward data.

Level 4: Reconciled attribution

At level four, evidence is matched to human time entries, projects, workstreams, and deliverables. People review exceptions rather than entering every fact manually.

An AI event that occurs during a recorded work period may attach to that entry automatically. If there is no confident match, it enters an orphan queue. A person can link it, edit the proposed attribution, confirm it, or dismiss it.

This creates an important separation:

  • Integrations capture observable facts.
  • The system proposes context.
  • A responsible person supplies judgment.

Typical capabilities

  • Automatic matching with visible provenance
  • Review queues for drafts and unlinked events
  • Deduplication across overlapping sources
  • Human confirmation and correction
  • Reporting by project, workstream, contribution, review state, tool, model, and cost

Main risk

Poor matching logic can make automation feel arbitrary. If people spend more time fixing the record than creating it, adoption will decline. The system should show why a match was proposed and make correction easy.

Level 5: Governed attribution

At the fifth level, attribution supports policy and business decisions across the organization.

The team can vary disclosure and evidence requirements by client, project, workstream, or risk level. Reports show only what the audience needs. Coverage thresholds prevent weak data from being presented as precise. Retention and deletion rules constrain what evidence is stored.

Typical capabilities

  • Approved tool and model policies
  • Client-specific disclosure postures
  • Required review or verification for selected workstreams
  • Coverage thresholds before quantitative reporting
  • Evidence retention and deletion controls
  • Internal and client-facing views from the same underlying record
  • Audit history for consequential changes
  • Billing rules based on actual agreements and recorded facts

Main risk

Governance can become bureaucracy. A mature system should make compliant behavior easier, not turn ordinary delivery into a form-filling exercise. Rules should be proportional to risk and tested with the people doing the work.

Continue improving

Review which fields are actually used in decisions. Remove capture that creates risk or effort without producing value. Maturity includes knowing what not to collect.

The capture methods are not the maturity levels

Ranking evidence sources from manual to automatic creates a misleading hierarchy.

Capture methodStrengthLimitation
Manual declarationCaptures intent, context, and human judgmentIncomplete and inconsistent without good workflow
Agent self-reportRich project and outcome contextCooperative and potentially incomplete
Provider usage APIAccurate model, token, cost, and timestamp dataOften weak project and outcome linkage
IDE or repository integrationDurable evidence close to coding workNarrow to specific tools and work types
Activity detectionBroad coverage with limited content captureShows presence, not exact contribution
Host-instrumented eventHigh-fidelity context from a controlled surfaceRequires each host to implement it

Mature attribution combines sources according to the decision and preserves provenance. A tool-presence signal can corroborate that AI was used. A human can describe why. A provider can supply cost. A repository can show the resulting change. No single source needs to pretend it knows everything.

Choose the level your decision requires

Not every team needs level five.

A small agency creating consistent client disclosures may get significant value from structured self-reporting. An enterprise with contractual AI restrictions may need connected evidence and governed policy. An engineering team allocating model cost may prioritize provider integration before client-facing disclosure.

Use the smallest system that can answer the real question reliably.

Ask:

  1. Which decision or obligation are we supporting?
  2. What is the minimum evidence that decision requires?
  3. Where does human judgment remain necessary?
  4. What content should we explicitly avoid collecting?
  5. How will we measure coverage and correct mistakes?

Those questions create a better roadmap than “integrate every AI tool.”

A practical 30-day starting plan

Week 1: Define

  • Name the first use case.
  • Define material AI involvement.
  • Choose involvement, contribution, and review labels.
  • Identify accountable roles.

Week 2: Record

  • Add the fields to an existing time entry, project workflow, or shared register.
  • Test the model on five recent deliverables.
  • Rewrite labels that people interpret differently.

Week 3: Review

  • Run a lightweight weekly review.
  • Identify missing records and unclear outcomes.
  • Decide what a client would see versus what stays internal.

Week 4: Connect

  • Select one high-value evidence source.
  • Define retention and privacy limits.
  • Measure whether the connection reduces effort or improves confidence.

The outcome after 30 days should not be a perfect dataset. It should be a shared attribution practice with evidence about what to automate next.

Better evidence, better decisions

AI attribution becomes valuable when it moves beyond a disclosure sentence and becomes part of the work record. That does not require a leap from memory to complete observability.

Start with language people understand. Attach it to real delivery. Add evidence where evidence changes a decision. Keep people in control of context and review. Govern the record in proportion to the risk.

Maturity is not the amount of data collected. It is the organization’s ability to explain how work happened, support that explanation with appropriate evidence, and act on it responsibly.

Related posts

AI did the work. Who gets the credit?

Human timesheets and AI usage logs each record only half the work. A useful attribution model keeps accountability human while making material AI contributions visible.

WhoWorked team8 min read

A practical model for AI attribution

A statement that says “AI was used” reveals almost nothing. Practical attribution records contribution, ownership, initiation, review, and evidence.

WhoWorked team9 min read

Start counting all the work.

30-day free trial. Bring the time history you already have.