Google’s bid for the operational archive of a failed airline is a useful anomaly: why should stale ERP records, old emails, source code and operational logs become valuable after the company itself has disappeared? The answer points to a different theory of data—less like a commodity and more like recorded organizational experience whose value depends on linkability, technological capability and future reuse.
The strange value of a dead company’s data
In August 2026, Google won a bankruptcy auction with a $10 million bid for a large body of de-identified enterprise data accumulated by Spirit Airlines before the carrier ceased operations. Reporting and bankruptcy materials describe a corpus that includes roughly 100 million emails, hundreds of millions of collaboration records, tens of millions of lines of source code, operational information, revenue and productivity data, audit material and large bodies of historical transaction and competitor-pricing observations. Customer databases and directly identifying customer information were represented as excluded from the proposed transfer.
As of 17 September, the transaction should still not be described as completed. The sale hearing has been adjourned to 30 September; six objections, a consumer privacy ombudsman report and a $12.5 million competing bid notice were on file by 11 September. The objections are not peripheral to the story. Employees, unions, suppliers and technology vendors have questioned whether de-identification is sufficient when records remain relationally linked and whether repositories may include intellectual property or confidential information belonging to third parties.
That procedural uncertainty actually makes the case more instructive. The important fact is not that Google has already acquired the archive. It is that sophisticated bidders assigned millions of dollars of value to the operational residue of a company whose operating business had disappeared.
The obvious explanation is that AI companies need data. It is also too weak. If the objective were simply to obtain more tokens, the archive of one failed airline would be an oddly cumbersome source. Public web text, synthetic data, licensed media, code repositories and commercial corpora exist at far larger scale. Spirit is interesting because its records are not merely text. They are traces of one economic organization interacting with the world over time.
The archive contains several partial descriptions of the same enterprise. Finance systems record transactions and resource states. Operational systems record events and constraints. Email and collaboration tools record negotiations, escalations and explanations. Source code records formalized rules. Tickets record failures and attempted remedies. External price observations record part of the environment in which decisions were made. The value is not located in any one source. It can emerge from the joins among them.
That turns a collection of records into something closer to a history of organizational behaviour: what the firm knew, what conditions it faced, what people and systems decided, which exceptions occurred, what action followed and what happened afterwards. The claim that such a dataset is literally the company would be absurd. The narrower claim is more serious: a sufficiently broad historical archive can contain part of the company’s accumulated learning curve in machine-readable form. This is where the old metaphor of data as the new oil starts to break.
Data was never really oil
The phrase was useful because it communicated scarcity, strategic control and the need for processing. Raw petroleum has little value to most users until it is extracted, transported, refined and distributed; raw data similarly requires collection, cleaning, interpretation and integration. The analogy also captured an early phase of the digital economy in which companies accumulated proprietary datasets and turned them into advertising, recommendation, fraud detection and optimization advantages.
But the analogy fails on the properties that matter most in the AI economy. Oil is rival. A barrel burned by one user cannot simultaneously be burned by another. Data is technologically non-rival: the same record can support accounting, forecasting, model evaluation, fraud detection and process analysis without being consumed by any one of those uses. Jones and Tonetti make this non-rivalry central to the economics of data because it can create increasing returns and unusual questions of ownership, access and market structure.
Oil is exhaustible. Data is reproducible at negligible marginal cost once collected, although legal rights and security controls can restrict that reproduction. A dataset can therefore be copied, recombined and reused across generations of technology.
Oil is comparatively fungible within a grade. A barrel of Brent is economically meaningful without knowing the identity of the refinery that will buy it. Enterprise data is radically context-dependent. Ten million disconnected rows from an ERP system may be nearly useless to an external observer if the identifiers, process semantics, organizational roles and surrounding events are unknown. A much smaller corpus with preserved provenance and relationships can be substantially more valuable.
Oil has an energy content that does not increase because a better refinery is invented. The usable value of an old dataset can increase sharply when a new analytical technology appears. Better entity resolution can reconnect records that were previously siloed. Better language models can extract semantics from emails and tickets that conventional analytics ignored. Multimodal models can connect images, voice, documents and telemetry. Agents can query several systems iteratively rather than requiring every possible integration to be built in advance. This last property is decisive. The economic value of data is partly a function of technologies that may not exist when the data is collected.
Diane Coyle and Luca Gamberi therefore propose treating data through a real-options lens: an organization may rationally preserve or invest in a dataset even when the eventual use case is uncertain, because ownership retains the right to exploit future analytical opportunities without obligating the firm to exercise them. In the AI era, this is more than an accounting subtlety. It explains why historical information can become strategically interesting after its original operational purpose has expired.
A better starting proposition is therefore not data is the new oil, but data can be recorded organizational experience with option value. That still needs refinement, because not all data creates the same option.
The value of data is relational, not additive
Executives are often shown data strategies as inventories: how many terabytes exist, how many sources are integrated, how many records have been catalogued, how much unstructured content has been indexed. These measures describe volume and accessibility, but not strategic value.
For enterprise AI, value often depends more on whether records can be related than on whether they can be counted.
Consider an exceptional procurement decision. An ERP database may record the purchase order, vendor, amount and approver. Email may contain the reason the normal supplier was bypassed. A risk system may contain a supplier alert. A ticket may show that a production line was already at risk. Source code or workflow configuration may encode the formal approval threshold. Telemetry may establish whether the intervention prevented an outage. Finance may reveal the eventual cost.
Each source alone answers a different question. Joined into one temporal episode, they can represent a decision under constraints and its outcome.
The same pattern appears in customer management. CRM data records an opportunity stage and eventual close. Meeting notes may explain why a discount was offered. Product telemetry records which features the customer actually used. Support tickets expose implementation friction. Finance establishes renewal economics. A model trained or evaluated on only the CRM learns a commercial state machine. A system that can connect the entire episode can begin to learn which contextual signals preceded a concession, when an exception was approved and whether the decision was economically successful.
The strategic unit of analysis is therefore not the row, document or database. It is the linked episode.
This has an important consequence: the value of two datasets can be greater than the sum of their separate values because one supplies the missing semantics of the other. Communications explain structured transactions; transactions validate claims made in communications. Code reveals which policies were executable; incidents reveal where those policies failed. External market data provides the counter-environment in which internal actions make sense.
That complementarity is one reason modern AI can revalue archives that conventional business intelligence could not justify integrating. The archive did not become richer. The cost of recovering relationships from it fell.
Why historical depth matters more than freshness
Much enterprise data governance implicitly assumes that old information loses value. For many operational purposes, this is correct. Yesterday’s inventory level is inferior to today’s inventory level. A former customer’s address may be useless. An obsolete product master can actively mislead.
But freshness and value are not the same variable.
If the objective is to reconstruct current state, old data depreciates quickly. If the objective is to learn how an organization behaves under different conditions, historical depth can increase value because it adds variation.
A decade of records can contain events that no current snapshot contains: recessions, supply failures, reorganizations, labor disputes, cyber incidents, failed product launches, pricing shocks, regulatory changes, acquisitions, software migrations, unusual customer behaviour and emergency workarounds. Routine cases are abundant; strategically important exceptions are rare. Historical depth expands the tail of the distribution.
This is particularly relevant to agentic systems. An operational agent that sees only normal cases may optimize the centre of a process while failing precisely where human expertise was historically most valuable. Rare events encode escalation logic, exceptions to policy and implicit boundaries of authority.
The paradox is that an old archive can be simultaneously stale as state and valuable as experience.
That distinction helps explain why Spirit is a better theoretical example than a live CRM database. A failed company’s archive can no longer tell Google how Spirit should operate tomorrow. It can, however, contain thousands of episodes showing how one airline responded to operational and commercial conditions over years. The company’s failure does not nullify every local capability recorded inside it. A firm can fail while still containing excellent maintenance routines, revenue-management practices, software components, operational playbooks and exception-handling mechanisms.
The hard problem is evaluation: distinguishing routines worth learning from those that contributed to failure.
A market is forming for different kinds of machine-legible experience
Spirit is unusual because the data became available through bankruptcy, but the broader market signal is visible elsewhere.
In 2024, Google entered into a licensing agreement with Reddit reportedly worth about $60 million annually, giving Google access to Reddit content for AI-related use. Reddit is valuable not merely because it contains text, but because it contains an enormous body of human questions, disagreements, niche expertise, judgments, feedback and social context. The corpus is a record of interaction rather than a static encyclopedia.
Google Cloud’s partnership with Stack Overflow provides a different type of asset. Google integrated Stack Overflow knowledge into its AI developer products while Stack Overflow adopted Google Cloud AI technology. Here the value is not simply source code or prose. It includes a highly structured relationship among technical questions, answers, community evaluation and domain-specific terminology. The metadata and interaction structure help distinguish plausible text from knowledge that has survived repeated scrutiny by practitioners.
OpenAI’s agreement with the Associated Press licensed part of an archive extending back to 1985, while its later partnership with the Financial Times included licensed content for model improvement and product experiences. The age of the archive is not incidental. Historical breadth exposes changes in institutions, vocabulary, actors and world conditions that cannot be reconstructed from a contemporary crawl alone.
Getty Images and NVIDIA demonstrate another model. Getty built generative products on a licensed creative corpus whose value includes not only images but curation, rights status and metadata. Getty explicitly positions the resulting model as commercially safer because the training corpus is controlled and licensed. In this case, provenance and permission become part of the data asset itself.
The inverse architecture is visible at Morgan Stanley. Rather than selling its proprietary research history to a model provider, the firm built AI systems that operate across its own corpus. AskResearchGPT was introduced to surface and synthesize an internal research estate producing more than 70,000 proprietary reports annually; another Morgan Stanley/OpenAI case describes an internal corpus of roughly 100,000 documents used by advisors. This is a strategically different use of the same technological principle. The model is made more useful by access to accumulated organizational knowledge, but control of the corpus remains with the enterprise.
These examples should not be collapsed into one category. Social discussion, legal content, journalism, visual media, corporate research and operational histories have different rights, quality characteristics and learning functions. Their common feature is narrower: AI increases the economic relevance of coherent corpora with history, metadata, provenance or internal relationships.
The emerging market is therefore not just a market for data. It is a market for different forms of machine-legible experience.
Synthetic data changes what is scarce
There is an apparent objection to this argument. If high-quality data becomes scarce, generative AI can increasingly manufacture more of it. Synthetic text, code, images, tabular records, simulated environments and artificial edge cases can be produced at a scale that would be impossible through human observation alone. If the supply of training data can itself be generated computationally, why should historical corporate archives become more valuable rather than less?
The answer requires separating data volume from empirical information. Synthetic data can be extremely valuable when the relevant structure is already sufficiently known. A simulator can generate millions of variations of a physical scenario. A model can produce additional programming exercises, paraphrases or labelled examples. A company can create artificial customer records for software testing without exposing real customers. Differentially private synthetic datasets can reproduce selected statistical properties of an underlying dataset while providing formal privacy guarantees that simple masking does not provide. Synthetic generation can therefore reduce the marginal cost of obtaining examples and can partly relax constraints imposed by privacy, labelling cost or rare-event frequency.
What synthetic generation cannot do by itself is create new observations of an unknown reality. A generated supplier failure is ultimately derived from assumptions about supplier failures already encoded in data, models or simulation rules. A synthetic negotiation can explore possible behaviours, but it does not reveal how an actual customer responded when millions of euros were at stake. A generated maintenance incident cannot discover an undocumented physical failure mode unless the generating system contains enough evidence or a sufficiently faithful model of the underlying physics to produce it.
This distinction is particularly important for enterprise data. Much of the value discussed in this article comes from historical contingencies: circumstances whose importance was not known before they occurred. A major outage, an unusual commercial dispute, a pandemic, an unexpected regulatory intervention, the failure of a strategic supplier or a sequence of apparently reasonable decisions that eventually destroyed value are informative precisely because reality produced combinations the organization did not necessarily anticipate.
Synthetic data can expand around observations of such events. It cannot retroactively manufacture the empirical fact that they happened, which signals were available beforehand, which people interpreted them correctly, which actions were actually authorized and what consequences followed.
Research on recursive use of generated data illustrates the deeper issue. Shumailov and colleagues found that indiscriminately training successive generative models on model-generated outputs can lead to model collapse, in which information about the original distribution is progressively distorted and, importantly, low-probability regions can disappear. Their result should not be interpreted as evidence that synthetic data is inherently harmful. Synthetic data can be deliberately generated, filtered, verified and mixed with empirical observations to improve training. The important point is epistemic: generated observations ultimately require some anchor outside the generative loop if the objective is to learn about the world rather than increasingly refined projections of an existing model.
This produces a paradox. The cheaper synthetic data becomes, the less strategically interesting generic data volume may become, while authentic observations of consequential reality can become relatively more valuable.
The distinction resembles the difference between simulation and measurement in engineering. Once a valid model exists, simulation can cheaply explore thousands of parameter combinations that could never all be tested physically. But the simulator has value because experiments and observations provided enough information to construct and validate it. Simulation multiplies the usefulness of empirical knowledge; it does not abolish the need for empirical knowledge.
For enterprise AI, the strongest architecture will therefore often combine both forms. Historical enterprise records provide grounding: real states, decisions, exceptions and outcomes. Synthetic generation can then expand the observed space, generate counterfactual cases, balance rare classes, create test scenarios and expose agents to situations that occurred too infrequently for direct learning. The original data supplies the anchor; synthetic data supplies controlled variation around it.
This also changes the interpretation of data scarcity. The scarce resource is increasingly unlikely to be tokens in the generic sense. Tokens can be manufactured. Examples can be generated. Even elaborate artificial environments can be simulated. What remains difficult to manufacture is a trustworthy observation linking what actually happened, under which conditions, because of which decision, and with which eventual outcome.
Spirit’s archive is interesting precisely for this reason. Google could generate billions of hypothetical airline emails or artificial booking transactions at negligible marginal cost compared with acquiring a historical corporate corpus. What it cannot simply generate is twenty years of observations produced by an actual airline interacting with actual customers, employees, competitors, aircraft, regulators and disruptions, together with whatever cross-system relationships survived in the archive.
Synthetic data therefore does not invalidate the argument that certain enterprise datasets are becoming strategically scarce. It makes the argument more precise. When artificial data becomes abundant, empirical provenance becomes scarce. When examples become cheap, observations become valuable. And when models can generate plausible organizational behaviour, records of what organizations actually did become increasingly important for distinguishing simulation from reality.
From data stock to organizational trajectory
The concept needed for enterprise strategy is trajectory. A dataset is a stock of observations. A trajectory links observations through time and action. It represents how a system moved from one state to another.
For an enterprise, a useful trajectory can connect an external condition, an internal state, an interpretation, a decision, an authorization, an action and an outcome. When many such trajectories are available, AI can do more than retrieve facts. It can compare cases, infer recurring decision patterns, identify exceptional paths, evaluate which signals historically mattered and support bounded simulations of alternative actions.
This is more than process mining, although process mining is an important precursor. Event logs can reconstruct the paths through which formal processes actually execute. Generative AI adds the possibility of incorporating unstructured explanations, negotiations, tickets and code that previously sat outside the event model. Knowledge graphs, entity resolution, vector retrieval, conventional machine learning, optimization and language models can all contribute. No single model needs to contain the whole representation.
The theoretical point is that organizational knowledge becomes more transferable when its traces can be converted into trajectories.
That does not make the organization transferable in full.
The firm is larger than its archive
A strong version of the Spirit thesis would claim that enough data permits an AI system to clone the company. There is no evidence for that claim, and several structural reasons to reject it:
- Archives are partially observable. They contain what was recorded, not everything participants perceived. A maintenance engineer may have reacted to a physical cue that never entered a system. A salesperson may have inferred hesitation in a conversation. A manager may have relied on trust built through years of interaction. The database records the action without necessarily containing the variable that caused it.
- Organizational competence includes tacit and embodied knowledge. Routines depend on people knowing whom to consult, which warning signs matter, how physical systems behave and when a formal procedure should be escalated. An archive can reveal that a particular person was repeatedly consulted without reconstructing why that person’s judgment was reliable.
- Historical data is mostly observational. It records the actions that were taken, not the outcomes of actions that were not taken. If an airline historically cancelled a flight under a particular combination of crew and weather constraints, the archive provides evidence about the chosen policy. It does not prove that cancellation was optimal. Causal evaluation requires experimentation, simulation, stronger structural assumptions or external evidence.
- Firms rely on complementary assets: physical infrastructure, licenses, capital, supplier relationships, trust, reputation, bargaining positions and legal authority. Data can describe how those assets were used; it cannot reproduce them by itself.
- Organizations are adaptive. Their capabilities are not static rules but results of repeated learning, changing incentives and path-dependent investment. An archive records one realized path through history. A new entrant begins from different initial conditions.
These limits matter, but they do not eliminate the economic implication. Competitive advantage does not require perfect non-replicability. It weakens when the cost or time required to imitate valuable capabilities falls.
A rival that can compress ten years of trial and error into two years of model-assisted learning has not cloned the incumbent. It has shortened the incumbent’s learning-curve advantage.
This is the more defensible meaning of capability extraction from enterprise data.
The new scarce resource is not data, but learning history
This reframing changes what executives should regard as strategic. The traditional data-management question is whether a dataset is accurate, current, governed and accessible. Those remain necessary. The AI-era question is whether the dataset encodes a difficult-to-recreate history of how the organization learned.
Some data is cheap to regenerate. Today’s public exchange rate can be fetched again tomorrow. A generic product description can be rewritten. A duplicated sensor measurement may add little once a physical phenomenon is already well characterized.
Other data cannot be regenerated without re-living the circumstances that produced it. A major supplier failure, a cyber incident, a product recall, an acquisition integration, a once-in-a-decade market shock, a complex negotiation or an emergency shutdown may never recur under comparable conditions. The surrounding communications, system states, actions and outcomes are therefore a unique observation of organizational adaptation.
This is why indiscriminate data retention is the wrong conclusion. Storage has cost; poor-quality archives create noise; privacy and regulatory obligations increase with retention; obsolete operational state can mislead; and preserving every trace creates security exposure.
The objective is not maximum data. It is maximum future learning value per unit of retained risk and cost.
That suggests a different classification scheme for enterprise data. Instead of asking only whether information is master data, transactional data, personal data or telemetry, executives can ask whether it has one or more of four properties:
- it captures a rare state or event that would be expensive or impossible to observe again;
- it links a decision to the information and authority available at the time;
- it records the outcome strongly enough to evaluate the decision later;
- it provides provenance or semantic context required to interpret other records.
Data with none of these properties may still be operationally necessary, but it is less likely to become a durable AI asset.
The option value of data rises when technology is moving quickly
Real-options thinking becomes especially useful when model capability is changing rapidly.
Suppose a company cannot currently extract much value from twenty years of maintenance reports because terminology is inconsistent, most files are unstructured and the reports refer to equipment identifiers that changed across several ERP migrations. A conventional ROI calculation may conclude that cleaning the archive is not justified.
That calculation assumes that the analytical frontier is stable. If entity-resolution systems, document understanding and agentic retrieval improve materially over the next three years, retaining the raw archive together with identifier mappings and provenance may preserve an option whose exercise cost later becomes much lower. Destroying the records is irreversible. Preserving them keeps future choices open.
This does not mean never delete data. The legal, privacy, cybersecurity and storage costs of retention are real. It means deletion decisions should account for technological uncertainty rather than valuing information only through currently funded use cases. Coyle and Gamberi’s real-options approach formalizes exactly this intuition: data may have value because future uses are uncertain, not despite that uncertainty.
The Spirit auction is a market signal consistent with this logic. The buyer is bidding not only on uses that can be specified today, but on the right to apply future techniques to a historical corpus that can never be recreated once destroyed.
Data also differs from oil because rights do not disappear when it moves
The economic analysis becomes incomplete if legal and organizational rights are treated as friction around an otherwise pure data asset.
The objections in the Spirit proceedings reveal that rights are part of the asset’s structure. A corporate archive is produced jointly by employees, customers, suppliers, software vendors and partners. A company may possess the storage system without possessing unrestricted rights to every semantic object inside it. Source repositories can contain third-party code. collaboration platforms can contain confidential partner information. Email can contain personal data. De-identification can remove direct identifiers while leaving relational patterns from which individuals or counterparties might be inferred.
This makes enterprise data less like a commodity and more like a bundle of conditional rights.
Getty’s generative-AI strategy shows the opposite side of the same phenomenon. It markets provenance and licensing discipline as part of the value proposition of the model itself. The rights status of the training data is not merely compliance overhead; it contributes to commercial usability.
For executives, this means that provenance, contractual permissions and retention rules can increase or destroy the option value of a corpus. A technically rich archive with unusable rights may have less strategic value than a smaller, well-governed corpus that can lawfully support model training, evaluation and retrieval.
AI turns data governance into capability governance
The conventional governance objective is to keep data accurate, secure, lawful and available to authorized users. In an agentic environment, another objective becomes necessary: controlling which parts of the firm’s accumulated learning can be exposed to external computational systems.
This is subtler than protecting confidential fields.
A customer name may be removable while the commercially valuable structure remains: ten years of negotiation history, discount patterns, escalation thresholds and contract concessions. Conversely, a model might need an account identifier to execute a narrow task without needing access to the history from which the firm’s bargaining strategy can be inferred.
The correct control objective is therefore not merely data minimization but semantic minimization: disclose the minimum business meaning required for the computation.
That has architectural consequences. Persistent enterprise memory, identity, policy, workflow state, evaluation history and tool authority do not have to reside inside the same model provider that supplies inference. If those objects remain enterprise-controlled, external models can compete for bounded tasks. If they migrate into one provider’s proprietary memory and orchestration layer, the organization may effectively transfer part of its machine-readable operating model even without transferring the underlying databases.
Morgan Stanley illustrates the alternative pattern: proprietary knowledge remains an enterprise asset and models are used to traverse it. This is strategically different from exporting the corpus in order to make an external platform progressively more knowledgeable about the firm.
The point is not that one architecture is always correct. It is that data strategy and model strategy can no longer be separated.
AI may create organization capital, not merely automate tasks
There is also a forward-looking implication. So far the discussion has treated historical data as a record of organization capital accumulated before AI. But AI can itself alter how organization capital is created.
An August 2026 NBER working paper by Babina, He and Jiang finds that recent firm-level AI investment is associated with productivity gains and traces those gains to the creation of organization capital: durable, firm-specific knowledge accumulated through learning-by-doing. The result is important because it suggests a two-way relationship.
Historical organization capital makes AI more useful because models can operate over accumulated firm-specific knowledge. At the same time, AI can accelerate the production and codification of new organization capital by making workflows observable, searchable, evaluable and easier to standardize.
This creates a strategic feedback loop. Firms that instrument decisions and outcomes well may learn faster from AI deployment. Faster learning produces richer trajectories. Richer trajectories improve future automation and decision support. The compounding asset is not a model weight file. It is the growing body of firm-specific, evaluated experience around the model.
That is much closer to the economics of organizational learning than to the economics of petroleum.
A more useful theory of enterprise data value
The Spirit case supports a compact theory. Enterprise data has no fixed intrinsic value. Its strategic value depends on at least six interacting properties:
- Uniqueness asks whether the observations can be regenerated cheaply. Public facts and routine measurements are less defensible than rare operational histories.
- Relational density asks whether records can be connected across entities, systems and time. The economic object is often the episode, not the row.
- Outcome visibility asks whether actions can be linked to consequences. Without outcomes, historical behaviour is easier to imitate than to evaluate.
- Technological accessibility asks whether available methods can extract the relevant structure at an acceptable cost. This variable changes rapidly with AI.
- Rights usability asks whether the corpus can actually be reused for training, retrieval, evaluation or product development.
- Complementarity asks what scarce assets must accompany the data: domain expertise, physical infrastructure, licenses, customer access, distribution, trust or execution authority.
These properties explain why two equally large datasets can have radically different values and why an old dataset can appreciate while a newer one becomes irrelevant.
They also explain why scale matters in a more precise sense than more data is better. Temporal scale captures rare events. Cross-system scale recovers missing context. Organizational scale exposes repeated routines and variants. Cross-enterprise scale, where legally and technically possible, helps distinguish firm-specific habits from more general regularities.
At that point, a corpus begins to resemble a dataset of organizational trajectories rather than a database of facts.
What Spirit should make an executive team ask
The Spirit auction is not proof that AI companies can reconstruct firms from their archives. It is evidence that historical enterprise data can retain significant option value after the operating company has disappeared.
That should provoke a different class of management questions. Which parts of our historical data record events that would be expensive to experience again? Can we reconstruct why significant decisions were made, under which authority and with which outcome? Do system migrations preserve the semantic identifiers required to connect old and new records? Are incident, project and commercial post-mortems linked to the operational evidence they discuss? Which external providers can see enough joined context to infer our processes rather than merely perform the task we contracted? Do our contracts preserve future reuse rights over the data we generate with suppliers and platforms? Which corpora would remain strategically useful if today’s applications disappeared?
These are not questions for the data office alone. They concern enterprise architecture, cybersecurity, legal rights, AI strategy and corporate memory at the same time.
Conclusion: data is the memory of experiments the firm has already paid for
Oil is valuable because nature spent geological time creating stored energy. Once extracted, it can be consumed. Enterprise data is valuable for almost the opposite reason. The organization spent money, time and managerial attention creating experiences: winning and losing customers, surviving incidents, negotiating contracts, changing software, reacting to competitors, making mistakes and correcting them. Data is valuable when it preserves enough of those experiences to let the organization—or another actor—learn from them again.
AI changes the economics because it reduces the cost of reading that memory. It can connect sources that were previously analytically separate, recover semantics from unstructured records, compare episodes across long periods and turn historical traces into inputs for retrieval, simulation, evaluation and bounded action. As those capabilities improve, the option value of some old corpora rises.
The correct conclusion is not that every company should hoard data indefinitely, nor that an archive can clone a firm. Much data remains redundant, stale, risky or legally unusable. Tacit knowledge, physical assets, relationships, incentives and institutional authority remain outside the archive.
The important shift is narrower and more consequential. Data should no longer be treated as a generic commodity whose strategic value comes mainly from volume. The most defensible enterprise data is the part that records a learning history competitors cannot cheaply reproduce: rare conditions, decisions, exceptions, actions and outcomes, linked with enough provenance to remain interpretable.
That is also why the Spirit case is so provocative. A dead airline’s current state is worthless. Its accumulated observations of how an airline operated are not necessarily worthless at all.
The AI-era version of the old slogan should therefore be almost the reverse of the original one.
Data is not the new oil. It is the memory of experiments the firm has already paid for—and AI is making more of those experiments readable.
AI Adoption, Productivity, and the Missing Middle
A review of Banca d'Italia QEF 1009 on Italy's AI adoption, firm productivity, and policy architecture
Process Mining in Manufacturing with Purchased Components and Variable Lead Times
A practical example of event-log analysis, object-centric modeling, lead-time diagnosis, and planning correction
Process Mining as Computational Process Intelligence
A formal and technical view of event-log analysis, object-centric process models, conformance checking, performance diagnosis, and enterprise applications
See also posts
When the User Becomes the Exploit
TerminalFix, WeWorm, and why security must survive the failure of its first boundary
Guerra profonda: recensione tecnica e guida all'approfondimento
Sovranità digitale, guerra algoritmica, AI e conflittualità ibrida nel libro di Arturo Di Corinto
When Formalization Became Industrial
Fermat's Last Theorem, Prove2Me, and the transition from proof abundance to formalization abundance
Programming Authority Is the Real PLC Security Boundary
Why network reachability, authentication, and controller programming must be treated as separate security states.
After Proof Abundance: Palomar and the New Infrastructure of Mathematical Trust
Formal verification, provenance, semantic fidelity, and institutional governance in machine-scale mathematics
The Industrialization of Mathematical Intelligence: Beyond Proof Abundance to Open Questions of Governance
Jacob Tsimerman, Timothy Gowers, and the changing roles of human judgment, agency, and mathematical culture
Back to top