The AI attribution maturity model
Teams can progress from ad hoc AI disclosure to governed work records without collecting every prompt or deploying every integration at once.

Teams often approach AI attribution as an instrumentation project.
They list every tool in use, investigate provider APIs, debate browser extensions, and ask whether they need to capture prompts. The project expands before the team has agreed on a more basic question: what decision should the attribution record support?
A consultancy trying to answer a client disclosure request has different evidence needs from an engineering organization measuring model cost. A regulated team proving human review needs different controls from an agency learning which workstreams are changing fastest.
A useful approach starts with shared terms, then improves the reliability of evidence as the team learns what its decisions require.
Maturity model
Better evidence, one level at a time
- Level 1
Ad hoc
Memory and scattered notes
- Level 2
Structured
Shared fields and definitions
- Level 3
Connected
Tool evidence supports declarations
- Level 4
Reconciled
Events link to real work
- Level 5
Governed
Rules match client and risk
Maturity has five dimensions
Before looking at levels, it helps to define what is becoming more mature.
Capture
How does the team learn that AI was materially involved? Through memory, a form, an integration, a provider feed, an agent event, or another signal?
Linkage
Can the AI contribution be connected to a person, project, workstream, and delivered outcome, or does it remain an isolated usage event?
Review
Can a responsible person confirm, correct, dismiss, or verify the record?
Governance
Are approved tools, disclosure rules, retention, privacy, and evidence requirements defined for the work?
Reporting
Can the organization use the record to answer a real question about delivery, billing, cost, client trust, capacity, or policy compliance?
A team can be advanced in one dimension and weak in another. Rich provider data with no project linkage is technically detailed but operationally immature. A simple manual disclosure tied to a client deliverable may be more useful.
Level 1: Ad hoc disclosure
At the first level, AI use is discussed but not recorded consistently.
One person adds a note to a document. Another mentions the model in a project channel. A third assumes that AI is just another tool and says nothing. When a client asks what happened, the team reconstructs the answer from memory.
Typical signals
- No shared definition of material AI use
- Disclosures vary by person and project
- Tool usage is visible only in separate vendor dashboards
- Human review is assumed rather than recorded
- Client questions trigger a manual investigation
Main risk
The team cannot distinguish “we believe this was reviewed” from “we have a record that it was reviewed.” The absence of structure creates both under-disclosure and overreaction.
Move to the next level
Agree on a minimal vocabulary: involvement level, contribution type, accountable owner, review status, project, workstream, and outcome.
Level 2: Structured self-reporting
At level two, the team records material AI involvement using a consistent form or template.
This is sometimes dismissed as manual or imperfect. It is also the fastest way to establish the operating model before investing in integrations.
Typical capabilities
- Shared definitions and examples
- A small set of involvement levels and contribution types
- Named ownership and review
- Records attached to projects and workstreams
- A client disclosure generated from the internal record
Main risk
Self-reporting is easy to forget and hard to enforce. People may interpret categories differently. The record says what someone remembers or declares, not necessarily everything that occurred.
Move to the next level
Identify the highest-value evidence gaps. Connect one or two important signals instead of trying to instrument every AI tool.
Level 3: Connected evidence
At level three, structured declarations are supported by signals from the tools where work occurs.
Evidence can arrive through several paths:
- A host application posts an event when a model is used.
- An AI agent reports a completed activity through an API or MCP tool.
- A provider usage API supplies model, token, cost, and timestamp data.
- An IDE or repository integration records relevant coding activity.
- A privacy-preserving extension records that an approved AI tool was active during a work period.
The goal is not to crown one source as universally best. Each source has a different balance of fidelity, coverage, deployment effort, and trust.
Typical capabilities
- Tool and model evidence supplements self-reporting
- Usage cost can be connected to a team or project
- Missing declarations can be surfaced for review
- Evidence provenance is recorded
- The team collects less content because metadata is sufficient for many decisions
Main risk
Connected data can create false confidence. A provider log proves that a model call occurred. It does not prove what business outcome resulted or whether a person reviewed the output. An agent that self-reports may produce rich context but remains a cooperative source.
Move to the next level
Reconcile events with the work record. Make uncertainty and missing linkage visible instead of silently discarding awkward data.
Level 4: Reconciled attribution
At level four, evidence is matched to human time entries, projects, workstreams, and deliverables. People review exceptions rather than entering every fact manually.
An AI event that occurs during a recorded work period may attach to that entry automatically. If there is no confident match, it enters an orphan queue. A person can link it, edit the proposed attribution, confirm it, or dismiss it.
This creates an important separation:
- Integrations capture observable facts.
- The system proposes context.
- A responsible person supplies judgment.
Typical capabilities
- Automatic matching with visible provenance
- Review queues for drafts and unlinked events
- Deduplication across overlapping sources
- Human confirmation and correction
- Reporting by project, workstream, contribution, review state, tool, model, and cost
Main risk
Poor matching logic can make automation feel arbitrary. If people spend more time fixing the record than creating it, adoption will decline. The system should show why a match was proposed and make correction easy.
Level 5: Governed attribution
At the fifth level, attribution supports policy and business decisions across the organization.
The team can vary disclosure and evidence requirements by client, project, workstream, or risk level. Reports show only what the audience needs. Coverage thresholds prevent weak data from being presented as precise. Retention and deletion rules constrain what evidence is stored.
Typical capabilities
- Approved tool and model policies
- Client-specific disclosure postures
- Required review or verification for selected workstreams
- Coverage thresholds before quantitative reporting
- Evidence retention and deletion controls
- Internal and client-facing views from the same underlying record
- Audit history for consequential changes
- Billing rules based on actual agreements and recorded facts
Main risk
Governance can become bureaucracy. A mature system should make compliant behavior easier, not turn ordinary delivery into a form-filling exercise. Rules should be proportional to risk and tested with the people doing the work.
Continue improving
Review which fields are actually used in decisions. Remove capture that creates risk or effort without producing value. Maturity includes knowing what not to collect.
The capture methods are not the maturity levels
Ranking evidence sources from manual to automatic creates a misleading hierarchy.
| Capture method | Strength | Limitation |
|---|---|---|
| Manual declaration | Captures intent, context, and human judgment | Incomplete and inconsistent without good workflow |
| Agent self-report | Rich project and outcome context | Cooperative and potentially incomplete |
| Provider usage API | Accurate model, token, cost, and timestamp data | Often weak project and outcome linkage |
| IDE or repository integration | Durable evidence close to coding work | Narrow to specific tools and work types |
| Activity detection | Broad coverage with limited content capture | Shows presence, not exact contribution |
| Host-instrumented event | High-fidelity context from a controlled surface | Requires each host to implement it |
Mature attribution combines sources according to the decision and preserves provenance. A tool-presence signal can corroborate that AI was used. A human can describe why. A provider can supply cost. A repository can show the resulting change. No single source needs to pretend it knows everything.
Choose the level your decision requires
Not every team needs level five.
A small agency creating consistent client disclosures may get significant value from structured self-reporting. An enterprise with contractual AI restrictions may need connected evidence and governed policy. An engineering team allocating model cost may prioritize provider integration before client-facing disclosure.
Use the smallest system that can answer the real question reliably.
Ask:
- Which decision or obligation are we supporting?
- What is the minimum evidence that decision requires?
- Where does human judgment remain necessary?
- What content should we explicitly avoid collecting?
- How will we measure coverage and correct mistakes?
Those questions create a better roadmap than “integrate every AI tool.”
A practical 30-day starting plan
Week 1: Define
- Name the first use case.
- Define material AI involvement.
- Choose involvement, contribution, and review labels.
- Identify accountable roles.
Week 2: Record
- Add the fields to an existing time entry, project workflow, or shared register.
- Test the model on five recent deliverables.
- Rewrite labels that people interpret differently.
Week 3: Review
- Run a lightweight weekly review.
- Identify missing records and unclear outcomes.
- Decide what a client would see versus what stays internal.
Week 4: Connect
- Select one high-value evidence source.
- Define retention and privacy limits.
- Measure whether the connection reduces effort or improves confidence.
The outcome after 30 days should not be a perfect dataset. It should be a shared attribution practice with evidence about what to automate next.
Better evidence, better decisions
AI attribution becomes valuable when it moves beyond a disclosure sentence and becomes part of the work record. That does not require a leap from memory to complete observability.
Start with language people understand. Attach it to real delivery. Add evidence where evidence changes a decision. Keep people in control of context and review. Govern the record in proportion to the risk.
Maturity is not the amount of data collected. It is the organization’s ability to explain how work happened, support that explanation with appropriate evidence, and act on it responsibly.