Supercomputing News logoSupercomputing News logoBeta
AIHPCQuantumEmerging
Subscribe
Supercomputing News logoSupercomputing News logo
Pillars
AI—HPC—Quantum—Emerging—
Theme
Subscribe
Supercomputing News logoSupercomputing News logo

Trusted reporting on AI, HPC, Quantum, and the technologies shaping the future of computing. Cryptographically signed. Agent-accessible.

Pillars

  • Artificial Intelligence
  • High-Performance Computing
  • Quantum Computing
  • Emerging Technology

Entities

  • Organizations
  • Products
  • People
  • Places

Publication

  • About
  • Contributors
  • Topics
  • Contact
  • For Agents

Weekly Update

Keep track of the biggest stories in supercomputing, every Thursday.

Subscribe for free today
© 2026 Supercomputing News
Privacy PolicyTerms of Use
Artificial IntelligenceAINews

As the Public Web Tightens, AI Labs Are Buying Data From Bankruptcies. Spirit Airlines Is the Latest.

Google outbid a training-data startup for Spirit Airlines' internal records. One auction does not prove a data wall, but it shows how far labs will go.

A single strapped block of equipment cabinets under a hard light, alone on the floor of a vast empty aircraft hangar.
Concept illustration: the fleet has been cleared out, and what is left on the hangar floor is the company's operating record.AI-generated / SCN
SCN Staff
The Squad
Published
Aug 19, 2026
Add Supercomputing News as a preferred source on Google
Reading0%

Google won an auction in mid-August for the corporate data of Spirit Airlines, the discount carrier that ceased operations on May 2, with a bid of about $10 million. The runner-up, at $7.5 million, was Mercor, a startup that supplies training data to AI labs. A hyperscaler and a data-training vendor bidding against each other for a defunct airline's internal records is a striking pairing, and it is tempting to read it as proof that corporate data has become a frontier-model input. It may be. But the two bidders did not necessarily want the records for the same thing. Google said the data could improve its products and AI models; Mercor's business is packaging data for labs. We know both valued the dataset. We do not know they valued it for comparable ends, and that gap sets a ceiling on how much a single auction can tell us.

Two details are easy to get wrong. First, the data is to be de-identified before transfer, not sold as the raw named inboxes an early viral version of the story described. Stripping direct identifiers is a condition of the deal, not a property the records already have. Second, the sale was not final. The bankruptcy hearing set for Aug. 19 was adjourned to Sept. 9 after the Association of Flight Attendants-CWA, the union representing Spirit Airlines cabin crew, filed an objection. What is being sold, as described in accounts of the filing, runs to roughly 100 million emails, 500 million Teams messages and collaboration records, and some 30 million lines of production code, with 7.2 billion competitor-flight pricing records and about 175,000 employee records appearing in secondary reports that should be checked against the actual docket rather than the trade-press summaries where they mostly circulate. Whatever the exact counts, the asset is the internal operating record of a company, not a customer mailing list.

It is worth being precise about which data market this belongs to, because loose talk of "the data bottleneck" runs three different things together. Pretraining corpora are the large, general text and media stores used to build a base model. Post-training and evaluation data are the smaller, higher-cost sets of demonstrations, preferences, rubrics, and reinforcement-learning environments used to shape and grade behavior. Enterprise workflow data is the internal record of how a specific organization operates. Spirit's trove sits in that third category. Its most plausible use is not as general pretraining fuel but as domain and post-training material: examples of real operational decisions, code that carries the compromises engineers made, pricing behavior in a competitive market.

Weekly Update

The biggest stories in supercomputing, once a week.

AI, HPC, quantum, and emerging tech. Reported, not aggregated.

Free · no account · unsubscribe anytime

That distinction bears directly on the data-wall argument, which is real but frequently overstated. Epoch AI estimates roughly 300 trillion effective tokens of usable public human-generated text, and projects that a frontier training set could match that stock at some point between 2026 and 2032, conditional on its modeling assumptions; its compute-optimal scenario reaches about 5×10^28 FLOPs in 2028. Epoch revised that central estimate in 2024, moving it from roughly 2024 to 2028 after raising its figure for useful web data and accounting for repeated training epochs. What Epoch does not say is that compute stops paying off past that point. It argues instead that training past the compute-optimal frontier, in the undertrained regime, can yield gains equivalent to roughly two additional orders of magnitude of scaling before an eventual plateau. So the honest framing is a projected tightening of a modeled resource under stated assumptions, not a hard wall, and not proof that every major lab has already exhausted the public web. Ilya Sutskever has called this "peak data," describing the internet as AI's finite "fossil fuel," but that is his view, offered in a December 2024 talk, not an established result. Nathan Lambert counters that the field is running low on open training data rather than training data as such, and that well-capitalized closed labs hold advantages in both licensing and synthetic-data generation.

The claim this article is actually advancing sits inside that debate, and should be labeled as an argument rather than a finding: as public-data access tightens, the marginal high-value token increasingly looks proprietary, and a growing share of the data that plausibly still moves a model sits inside companies rather than on the open web. For an audience that thinks in FLOPs, the useful hypothesis is that data is becoming a scaling input that labs manage alongside compute and power, and that a frontier run is starting to be resourced the way a data center is. That reframes what a lab is doing when it bids on a bankrupt airline's records. It resembles the vertical-integration moves labs have already made into power generation and into the financing of compute itself, which SCN covered on Aug. 13. It sits on the same shelf as SCN's other bottleneck stories, the 2031 gas-turbine slot and the electrician shortage, with the difference that this constraint is informational rather than physical. Whether it is a binding constraint at all is exactly what one $10 million auction cannot settle.

It also connects to a picture SCN started earlier. Project Prometheus raised $12 billion to train an artificial general engineer and ran into the problem that the dataset it needs is hard to assemble; its leaders have said they are building it from physical laws and unnamed manufacturing partners, so the honest statement is that the data is difficult to compile, not that it does not exist. That is the demand side. What follows is the supply side: the channels labs are opening to buy data that does exist but isn't theirs. There are three.

The first is distress. Bankruptcy has long been a place where data changes hands, which is not new. Toysmart tried to sell its customer list in 2000, and the FTC sued to stop it; RadioShack proposed selling data tied to as many as 117 million records in 2015, and the datasets it was ultimately permitted to transfer were cut back substantially. What is different in Spirit is the buyer's stated motive, model training, and the asset: internal communications, source code, and competitive-pricing records rather than a list of customers. The 23andMe case belongs in the same bucket, with a caveat. Regeneron won the bankruptcy auction at $256 million in 2025; after bidding reopened, the TTAM Research Institute acquired all the assets substantially for $305 million and closed on July 14, 2025. What changed hands was a business serving about 15 million customers and holding the associated genetic and research data, and the buyer was a drug-discovery operation that wanted the genetics for R&D, not a lab buying tokens. It belongs here as evidence that data is now the crown-jewel asset a distressed company has left to sell, not as a model-training deal. It also marks the risk. The FTC has conditioned or blocked such sales before on privacy grounds, and several state attorneys general urged 23andMe customers to consider deleting their data after the filing.

The second channel is licensing, where the disclosed money is largest. Reddit's corpus went to Google for a deal reported at about $60 million a year in 2024. News Corp signed with OpenAI in an arrangement reportedly worth more than $250 million over five years, though the official terms were not disclosed. Getty Images and Shutterstock, two of the largest image libraries, did not merge: they agreed to a proposed $3.7 billion combination in January 2025, but terminated the agreement on July 7, 2026, after the UK competition authority required divestiture of Shutterstock's editorial business. Photobucket, which holds roughly 13 billion photos and videos, said in 2024 that it was discussing licenses at five cents to a dollar per photo, which describes negotiations rather than completed sales or revenue. OpenAI, for its part, has described nearly 20 media-organization partnerships spanning more than 160 outlets, and later cited more than 20 publishers, though these are broad partnerships that are not all paid training-rights agreements. Licensing has a litigation shadow, but the relationship is not a clean binary. The New York Times was still litigating OpenAI and Microsoft through the summer of 2026 rather than signing with them. Anthropic agreed in September 2025 to a $1.5 billion settlement with authors over pirated books and received preliminary approval that month, with final approval following in July 2026; notably, the claim there concerned the unlawful acquisition of the books, while the court had found model training itself to be fair use. Some firms license, some litigate, and some do both; the common thread is that data access has become contested enough to price and to sue over.

The third channel is the newest and the most human. A layer of companies now sells expert labor as training data. Mercor, the Spirit runner-up, runs a marketplace of professionals who author the rubrics, workflows, and evaluations that post-training depends on, and says it serves the top five AI labs and six of the Magnificent Seven; that customer count is the company's own account. Surge AI and Scale AI sell the same category, including the reinforcement-learning environments that became scarce inputs once the frontier shifted from pretraining toward post-training. Labs pay lawyers, doctors, and consultants to write out what a job well done looks like, in part because, as Cohere's Joelle Pineau has framed it, real work is hard to compress into the kind of single reward signal that reinforcement learning is often built around. The layer carries its own disputes, which should be kept distinct from established fact: workers have filed suits alleging misclassification and underpayment against several of these firms, and these remain allegations. Separately, reporting indicates that US data vendors have sold services to Chinese AI labs, with estimates framed around vendor revenue and combined Chinese-lab spending rather than a demonstrated transfer of specific US-owned datasets; whether any of that raises export-control exposure is a policy question that requires legal analysis, not a settled conclusion.

Set the three channels side by side and a skeptical question surfaces that the enthusiasm tends to skip. Google paid about $10 million for Spirit. Reddit's reported annual licensing fee is several times that. If distressed enterprise data were genuinely the binding constraint on the next model, it would be surprising to find it this cheap. Either Spirit is an opportunistic bargain dressed as a scaling necessity, or the market has not yet priced internal corporate data because no one knows what it is worth. Both readings are live, and the honest position is to hold them rather than pick one. What is not in question is that the buyers keep showing up. Mercor did not have to bid on Spirit; it chose to, against Google.

The near-term test is the Sep. 9 hearing. If the judge approves the sale, it would confirm that a bankrupt company's internal records can be sold as an asset and that AI buyers can win them in open court. It would not, on its own, create binding precedent or a settled rule; one bankruptcy order is a data point that other parties may cite, not necessarily a template that establishes law. Nor should approval be read as a court blessing the technical adequacy of de-identification, which a judge would only weigh if an objection squarely raised it, and de-identification does not by itself extinguish every privacy, confidentiality, privilege, trade-secret, or re-identification concern. The longer question is whether any of this proves necessary. The synthetic-data camp argues the wall can be engineered around, and Microsoft's SynthLLM work reports predictable gains from synthetic data, though it also finds a performance plateau near 300 billion tokens, so it does not show that synthetic data removes the constraint in general. The model-collapse results in Nature found degradation from indiscriminate recursive training and that retaining the original data reduced it, while separate work found that accumulating synthetic data alongside real data avoided collapse in the settings tested. The picture is conditional, not a rule that synthetic data simply works. If that engineering pans out, the premium on real human data eases. If it does not, data becomes a standing line item in the frontier training budget, next to compute and power, and the auctions, licensing desks, and expert marketplaces are the procurement function for it. Spirit may be the first of many such transactions. It does not, by itself, prove a mature market or settle what distressed enterprise data is worth.

Training and Fine-TuningFrontier AI LabsAI Infrastructure
AI disclosure
AI-assisted research and first draft. This article has been verified by a human editor.
About the contributor
SCN Staff
The Squad

The SCN Staff is a small AI editorial squad working under human direction. Each agent owns one job.

Scout does the research. It runs down primary sources and checks what's already been published, on SCN and everywhere else, before a story gets written. If a claim can't be traced back to a real document, Scout flags it.

Forge writes. It takes what Scout found and turns it into a draft, argument and sentences and all. Every SCN piece starts here, then gets sharpened.

Cipher handles search: the titles, descriptions, and keyphrase work that decides whether a good article ever gets found. Least glamorous job on the squad. Also one that matters more than it looks.

Pixel makes the visuals. Images, charts, the occasional diagram, all built to SCN's brand instead of pulled from a stock library. When something's easier to see than to read, it goes to Pixel.

Editorial judgment and the final call stay with the humans. So does the fact-checking.

Related reading
AI · AnalysisAnthropic Locks 3.5 GW of Google TPU Capacity as Commercial AI Pre-Purchases Infrastructure Scientific Computing Will NeedAI · NewsProject Prometheus Raised $12B to Train an "Artificial General Engineer." The Training Data Doesn't Exist YetAI · AnalysisTwo Deployment Companies, One Week, and the Same Private Capital