The first giant AI copyright settlement has reached the least glamorous part of the machine: figuring out who actually gets paid. That sounds like accounting. It is not. It is a stress test for the ownership infrastructure underneath the publishing industry, and the infrastructure is coughing up dust.

During the week ending September 4, the administrator for the $1.5 billion Bartz v. Anthropic settlement sent reconciliation notices to claimants. The notices showed who else had filed against the same title and what share each claimant requested. For many works, confirmation should be routine. For others, the portal surfaced disagreements among authors, publishers, co-authors, and, according to fresh reporting, some literary agencies. A settlement designed to compensate rightsholders has become a live audit of whether the industry can identify them.

The new notices do not prove a coordinated land grab. The Authors Guild says certain publishers told the administrator they mistakenly selected a 100 percent allocation instead of the default split, and that those selections were being corrected. TechCrunch reports complaints involving publishers that allegedly claimed reverted works or requested the full payment, along with reports of agencies seeking a percentage. Those claims have not been adjudicated. Bad records, confusing forms, aggressive contract readings, and genuine ownership conflicts can all produce the same ugly portal screen.

That qualification matters because outrage travels faster than contract law. A publisher appearing beside a book does not automatically mean the publisher is stealing from its author. An author believing rights reverted does not automatically settle the date or scope of that reversion. An agent taking a commission on normal book income does not automatically establish a legal or beneficial ownership interest in a copyright recovery. Each answer can depend on the agreement, the chain of title, the right allegedly infringed, and who held it when the infringement occurred.

The official process starts from a cleaner rule. For non-education works, the default option sends 50 percent of the per-work payment to valid author claimants and 50 percent to valid publisher claimants, divided within each group when multiple parties qualify. Claimants can propose an alternative allocation, but the percentages must total 100 percent. They can also challenge whether another listed party is a rightsholder and upload contracts, reversion letters, or other supporting documents.

The estimated headline payment is approximately $3,000 per eligible work before costs and fees. That is not a promise that every writer receives a $3,000 check. One work can have multiple authors, publishers, editions, contractual arrangements, and competing claims. The settlement covers 482,460 works, according to the court's final approval order. By July, valid claims had reached at least 91.3 percent of those works. Even a small error rate at that scale becomes thousands of human disputes.

The court anticipated this. In November 2025, it appointed Theodore K. Cheng as a neutral Special Master for ownership and allocation conflicts that claimants and the administrator could not resolve. The Authors Guild describes a sequence that begins with a 30-day window for co-claimants to reach agreement, followed by administrator-assisted resolution, then referral to the Special Master. His decisions are final and binding without an appeal. Payments tied to a disputed work wait until the disagreement is resolved.

The process is legally serious and operationally brutal. The final approval order requires submissions in the Special Master process, including publishing agreements, to remain confidential and under seal. That protects sensitive contracts, but it also means the public will not receive a clean dataset explaining why one author received everything, another split the payment, and a third lost a claim. The system resolves individual cases without automatically creating the shared rights map the market still needs.

The underlying data problem starts earlier. Anthropic downloaded roughly seven million files from LibGen and PiLiMi and produced them during litigation. Plaintiffs used those files to construct the Works List that defines the class. The settlement's own search guidance says identifiers and metadata came from the downloaded files, copyright records, Bowker Books Data, ISBNdb, and manual review where possible. It also warns that records can disagree because editions, formats, titles, and contributor lists change.

That is a forensic reconstruction, not a modern rights system. An ISBN identifies a publication product. A copyright registration identifies a registered work or claim. A publisher database may reflect an edition. A contract defines a bundle of rights that can be licensed, assigned, reverted, terminated, split by territory, divided by format, or shared among contributors. Those objects overlap, but they are not interchangeable. Dump them into one table and the rows may look precise while the ownership conclusion remains wrong.

Rights reversion makes the weakness visible. The Authors Guild argues that an author whose rights reverted before the settlement's August 10, 2022 download date may be entitled to the full payment, depending on the contract. A later reversion can produce a different result because the publisher may have held the relevant reproduction right when the alleged infringement happened. The difference between 100 percent and 50 percent may therefore turn on one letter, one clause, and one date that never entered a shared database.

Educational publishing adds another layer. Trade contracts often license specified rights while leaving other rights with the author. Educational agreements can assign much more. The Authors Guild says some educational publishers are applying ordinary royalty percentages to settlement proceeds when contracts do not expressly address infringement recoveries. An author may read the payment as compensation for a violation of personal rights. A publisher may read it as proceeds governed by the commercial agreement. The portal cannot settle that philosophical fight. It can only force the parties to document their legal one.

The agent question is even messier. TechCrunch reports complaints about agencies filing for percentages, while critics argue that agents are not rightsholders in the books they sell. That may be true in a typical representation agreement, but the responsible conclusion is narrower: a commission clause and a copyright ownership interest are different things, and nobody should confuse them without reading the contract. The settlement portal records requested allocations. It does not certify every request as valid merely because someone submitted it.

This is where the AI story gets larger than Anthropic. Every proposed licensing market for training data assumes a reliable answer to four questions: what is the work, who controls the relevant use, what permission exists, and where should the money go? Publishing can often answer those questions through lawyers, royalty departments, agents, collection societies, and institutional memory. It struggles to answer them instantly, consistently, and in a format that software can execute across millions of works.

That gap creates a category. The industry does not need another dashboard that scrapes book metadata and calls itself a rights platform. It needs durable, machine-readable records that connect works, editions, registrations, contracts, territories, formats, dates, reversions, contributors, and payment instructions. Each record needs provenance. Each change needs an audit trail. Conflicts need a governance process. Sensitive contracts need privacy controls. The useful product is not a prettier catalog. It is an ownership graph with evidence.

Model developers need the same layer for a different reason. A license is only as good as the licensor's authority. Paying the wrong party does not magically erase exposure from the party that actually controls the right. At industrial scale, one-to-one diligence becomes impossible. Developers will need standardized rights assertions, warranties, update mechanisms, conflict flags, and payment routing that can survive acquisitions, reversions, estates, and new editions. Otherwise licensing becomes expensive theater performed on top of bad data.

Creators should want this infrastructure too, but not if it becomes another black box controlled by the largest publishers or model labs. Authors need visibility into which works were licensed, which rights were asserted, what rate applied, how deductions were calculated, and where disputes can be challenged. A machine-readable registry without human appeal is just automated opacity. The point is not to remove judgment. It is to stop wasting judgment on facts the system should already know.

The settlement administrator is doing something closer to emergency data integration than ordinary check processing. The court says notice was sent to 594,945 potential class members associated with 482,374 works, using sources that included author groups, more than 170 publishers, copyright records, commercial databases, and targeted searches. That effort reached rightsholders associated with 99.5 percent of the Works List without returned notice. It is an impressive campaign. It is also a ridiculous way to rediscover an industry's ownership map every time a new technology creates value.

The economics make the lesson harder to ignore. The fund is non-reversionary, meaning unused money does not simply return to Anthropic. The court awarded class counsel about $101.56 million, nearly 6.8 percent of the fund, and held back 10 percent of that fee pending post-distribution accounting. Administrative and future dispute costs also consume resources. None of this means the process is unnecessary. It means bad rights infrastructure has a price, and the price shows up before a creator receives anything.

The settlement itself remains narrower than the rhetoric around it. The court approved compensation tied to Anthropic's acquisition and copying of books from pirate sources. The fair-use ruling treated model training differently from acquiring pirated copies, and the final agreement does not resolve every copyright question about AI outputs or future conduct. The current allocation fight is therefore not a referendum on whether generative AI should exist. It is a test of whether compensation can find the people legally entitled to it.

The strongest response is not another round of slogans about creators versus machines. It is to build the missing transaction layer. Publishers should digitize rights chains and reversions. Agents should distinguish commission claims from ownership claims. Authors should keep contracts and reversion records accessible. Model companies should demand auditable authority before treating a license as clean. Registries should interoperate instead of pretending one identifier can represent every relevant fact.

Anthropic's settlement put a dollar value on one historical collision between AI development and pirated books. The reconciliation process is revealing the next bottleneck. We can negotiate billions in principle, then lose months deciding which human owns which percentage because the evidence lives in old PDFs, disconnected databases, and somebody's filing cabinet. That is not merely administrative friction. It is an invitation to build the financial and legal infrastructure that an intelligence economy will require.

The money will eventually move. The deeper signal is what had to happen before it could. The first major AI copyright settlement did not find a mature licensing market waiting for it. It found a stack of contracts, a fragmented identity problem, and a court-appointed referee. Anyone serious about the future of paid data should stop treating rights metadata as paperwork. It is the product.

LaunchPad positionAI licensing cannot scale on contracts trapped in filing cabinets and ownership data scattered across incompatible registries. The companies that build machine-readable rights records, provenance, and payment routing will own a critical layer of the information economy.
Reporting standard

This report draws on the linked primary sources and reputable reporting. Company statements are treated as claims until independently demonstrated.