(A follow-up to "When the Tool Does the Work: How Do You Still Measure the Human?")
Last time I argued that AI changed the line we have to measure against. A competent result is no longer enough to prove human skill, because a machine can now produce competent results on its own. The human signal is the margin above that machine baseline.
But a margin by itself is not yet something a market can trust.
Before an owner delegates a deck, before an employer pays a premium, before anyone treats that margin as evidence of skill, one harder question has to be answered:
Whose margin is it — and where are the receipts?
Suppose a player has a 62% win rate against a machine baseline of 50%. That twelve-point gap looks like the human signal we want.
Maybe it is. But the number alone does not tell us enough.
Was the same person making the decisions? Were the matches played under comparable conditions? Did the player have a much stronger deck? Was a tool assisting the choices? Was the result a short lucky streak, or part of a durable pattern?
The uncomfortable truth is that a score is also an output — and AI is very good at producing outputs that look convincing.
So the same rule applies twice. First, the raw result needs to clear the machine's bar. Then the claimed human contribution needs evidence that survives inspection.
The human margin needs a chain of custody.
This distinction is going to matter everywhere.
A password proves that someone can open an account. A wallet signature proves control of a key. A biometric may prove that a person was present for a check. None of those facts, by themselves, prove who performed the cognition behind the work that happened next.
You can hand your keys to a script. You can authenticate and then let an agent take over. You can be a real person operating an account whose actual output is almost entirely automated.
So human provenance is not simply:
Is there a human attached to this identity?
It is:
What evidence shows the human contribution that mattered in this particular event?
Identity tells us who can open the door. Provenance tells us what happened after the door opened.
The receipts are not a webcam pointed at someone for eight hours. They are not a demand to expose every private thought or every click.
They are the smallest reliable record of the conditions and choices that explain the result:
That creates a trail from the result back to the conditions under which it was earned.
Think of a receipt at a store. The number charged to your card is not enough. The receipt names the items, the quantities, the price of each, the time, and the merchant. It lets another person reconstruct why the total is what it is.
A human-skill signal needs the same property. Not merely a number, but a number whose causes can be followed backward.
Even perfect receipts from one event cannot separate skill from luck.
Anyone can make one brilliant call. Anyone can win one impossible match. Anyone can also make one terrible decision that says nothing about what they are normally capable of.
The signal becomes useful when the same kind of human margin appears repeatedly:
Now we are no longer looking at a screenshot. We are looking at a trajectory.
Did the player adapt? Did the advantage survive when the deck changed? Did judgment improve after failure? Did performance collapse outside one memorized pattern, or travel across unfamiliar conditions?
Time matters because history is difficult to manufacture retroactively. A new identity can copy a profile picture and generate a polished résumé in an afternoon. It cannot instantly acquire years of coherent decisions made under real constraints against real counterparties.
That accumulated history is provenance capital.
This is where Splinterlands becomes much more than an example.
Each battle exposes a structured contest: the players, the rulesets, the mana limit, the available elements, the deployed cards, the resource imbalance, the outcome, and a replayable record of what happened. The conditions keep changing, so a player cannot prove much by repeating one fixed answer forever.
That gives us the ingredients for a real chain of custody:
The win is not the receipt. The battle record is the receipt.
The win rate is not the human. The repeatable margin above what the machine, the cards, and the conditions already explain is the human signal we are trying to preserve.
Return to the scholar system.
Today an owner often delegates through reputation: a Discord relationship, a recommendation, a screenshot, or a small circle of prior trust. That works while the market is small. It does not become a global labor market until the evidence can travel farther than the relationship.
Imagine instead that a scholar carries a portable record showing:
Now the owner does not have to trust a stranger's biography. The owner can inspect the work.
And the scholar is no longer trapped inside one guild's private reputation system. Their demonstrated ability belongs to them. It can move with them to the next owner, the next contract, and eventually the next kind of work.
Capital can finally find talent through evidence.
The temptation, of course, is to turn all of this into a single score.
That would be convenient. It would also throw away most of what makes the evidence valuable.
Two players can reach the same rating by completely different paths. One may dominate with expensive resources. Another may consistently outperform weaker tools. One may be brilliant in stable conditions and collapse when the rules change. Another may lose more often overall but repeatedly make the difference in close, ambiguous matches.
Those are not the same worker.
A useful record has to preserve the components: pressure, asymmetry, coordination, adaptation, judgment, and the evidence behind each one. A summary can sit on top. It cannot replace the structure underneath.
The market should be able to ask not only, "How good is this person?" but, "Good at what, under which conditions, and according to which evidence?"
A trustworthy system also has to know when it does not know.
Missing evidence is not a zero. A new player is not unskilled merely because the history is short. A private or unavailable record is not proof of automation. A strange result is not proof of fraud.
If the coverage is too thin, the answer has to be: not enough evidence yet.
That sounds less exciting than a bright green badge declaring "verified human." It is also the only foundation a serious market can survive on. The first time a system confidently invents certainty, every honest participant begins paying for its mistake.
The goal is not an infallible human detector. The goal is a bounded claim with visible evidence, visible limits, and a trail another person can audit.
The same problem follows us into every AI-mediated market.
When AI writes competent code, the valuable developer is the one whose judgment improves the result above what the tools already produce — and whose decisions, reviews, recoveries, and outcomes leave evidence.
When AI creates competent media, the market for human authorship needs more than a byline.
When AI advises a community, legitimacy depends on knowing where accountable human judgment entered the decision.
When researchers need actual human behavior, synthetic responses are not cheaper data. They are contaminated data.
Not every market will care. If replacing the human changes nothing about the good, human provenance may carry no premium at all.
But wherever the human contribution is part of what the buyer, community, or institution actually values, the receipts become infrastructure.
The chain now looks like this:
The result tells us what happened. The machine baseline tells us what was already available to everyone. The margin tells us what may have been added. The receipts show how it was earned. The history shows whether it belongs to a durable human trajectory.
That is enough to turn performance into something a market can begin to trust.
But it opens a much larger question.
What happens when those receipts no longer come from only one game? What happens when the same person leaves evidence through competition, work, creation, relationships, and consequential choices — and that evidence begins accumulating into one continuous economic history?
At that point we are no longer building a better leaderboard.
We are building something much closer to a new kind of identity.
That's where the story goes next. 🛡️