Abstract

The governance of artificial intelligence is overwhelmingly theorized through two institutional frames. In the market frame, the data, models, and compute that constitute the AI stack are private goods exchanged under property and contract; in the state frame, a regulator imposes rules from above. A third possibility, the collective and self-organized stewardship of AI-relevant resources by the communities that produce and depend on them, remains comparatively under-theorized, even as it proliferates in practice through data trusts and cooperatives, federated learning consortia, public compute initiatives, open-weight model collaborations, and community data sovereignty regimes. This article argues that these arrangements form a coherent institutional family, which we call commons-governed artificial intelligence, and that the analytic vocabulary developed by Elinor Ostrom and her successors for common-pool and knowledge commons is the right backbone for classifying them. We contribute a two-dimensional taxonomy whose first axis is the resource layer of the AI stack held in common, distinguishing data, compute, models, knowledge and evaluation, and energy, and whose second axis is the governance function performed, derived from Ostrom’s design principles of boundary definition, appropriation and provision congruence, collective choice, monitoring, graduated sanctioning, conflict resolution, recognition of the right to organize, and nested polycentric scaling. We populate the taxonomy by examining the published evidence layer by layer, locate ten recurrent institutional arch...

Keywords: AI governance; commons; Ostrom; common-pool resources; data trusts; data cooperatives; federated learning; compute governance; open-weight models; sustainable AI; polycentric governance.

1 Introduction

Artificial intelligence is now built from a stack of resources whose scale and concentration are without precedent in the history of computing. Frontier models are trained on corpora assembled from the textual and visual output of much of humanity, on clusters of specialized accelerators available to only a handful of firms and states, and at an energy cost that has become a measurable fraction of global electricity demand (International Energy Agency, 2025) . The question of how this stack should be governed has accordingly moved from a specialist concern to a central problem of contemporary political economy. Yet the dominant ways of posing that question presuppose one of two institutional answers. The first is the market answer, in which data, model weights, and compute are private goods, allocated through property and contract, and governed by the firms that own them. The second is the state answer, in which a public regulator imposes binding rules from above, as in the European Union’s Artificial Intelligence Act (European Parliament and Council of the European Union, 2024) , the risk-management framework of the United States National Institute of Standards and Technology (National Institute of Standards and Technology, 2023) , or the intergovernmental principles of the Organisation for Economic Co-operation and Development (Organisation for Economic Co-operation and Development, [2019](https://arxiv.org/html/2606.15466v1#bi...

The claim that there exists a coherent third way is not a rhetorical flourish but the central empirical finding of the commons tradition in institutional analysis. Hardin (1968) argued that a resource held in common is doomed to overexploitation, because each appropriator captures the full benefit of an additional withdrawal while bearing only a fraction of its cost, so that, in his phrase, freedom in a commons brings ruin to all. The policy corollary, that only privatization or external coercion can avert collapse, is precisely the dichotomy the market and state frames inherit. The decisive rebuttal is the body of work for which Ostrom (1990) received the Nobel Memorial Prize in Economic Sciences. By distinguishing open-access regimes, which were Hardin’s actual model, from common-pool resources governed by bounded communities under self-determined rules, and by documenting durable self-governing institutions across irrigation systems, fisheries, forests, and groundwater basins, Ostrom established both theoretically and empirically that the space between markets and states is institutionally occupied. Her account of polycentric governance, the thesis that complex resource systems are stewarded effectively by multiple overlapping centers of authority rather than by a single monocentric one, gives the third way its structure (Ostrom, 2010) . The organizing wager of this article is that the resources of the AI stack, training corpora,...

That wager is supported by a rapidly growing but fragmented practice. On the data layer, bottom-up data trusts vest fiduciary control of pooled data in a trustee bound to the interests of the data subjects (Delacroix and Lawrence, 2019) , data cooperatives place curation and analytics under member ownership on the model of the credit union (Hardjono and Pentland, 2019) , and community regimes such as the CARE principles for Indigenous data governance assert collective rather than merely individual control over data (Carroll et al. , 2020) . On the compute layer, the United States National AI Research Resource and the European EuroHPC initiative pool publicly funded computational capacity for broad access (National Artificial Intelligence Research Resource Task Force, 2023; European High Performance Computing Joint Undertaking, 2024) , and federated learning makes it possible to train a shared model over decentralized data without pooling the raw data itself (McMahan et al. , 2017; Kairouz et al. , 2021) . On the model layer, collaborations such as the BigScience workshop have produced openly licensed large language models on publicly granted supercomputers (BigScience Workshop et al. , [2022](https://arxiv.org/html/2606.15466v1#bib.bib58...

The two literatures that come closest each stop short of the category we propose. The AI-governance literature is organized around principles, risks, and regulatory compliance rather than around institutions. The most cited map of the normative terrain, the meta-analysis of Jobin et al. (2019) , documents convergence on five ethical principles across eighty-four guideline documents, and the synthesis of Floridi et al. (2018) consolidates them into a unified ethical framework; but neither classifies the institutional forms through which such principles might be enacted. Mittelstadt (2019) argues precisely that principles alone cannot guarantee ethical AI, because the principlist analogy with bioethics fails in the absence of common aims, fiduciary duties, professional histories, and accountability mechanisms. The most comprehensive recent stocktake, the AI Risk Repository of Slattery et al. (2024) , is a database and taxonomy of risks, deliberately agnostic about institutional remedies. The data-governance literature, by contrast, is institutionally rich but was largely written before the foundation-model era and is organized around personal data rather than AI training assets. Its proximate taxonomy, the four emerging models of Micheli et al. (2020) , namely data-sharing pools, data cooperatives, public data trusts, and personal data...

We close that gap with a two-dimensional taxonomy. The first axis is the resource layer of the AI stack that is held in common. We distinguish five layers, data, compute, models and weights, knowledge and evaluation, and energy, and we argue that each satisfies the structural criteria for treatment as a commons, in particular that each is subtractable in some governance-relevant dimension even where its informational content is non-rival (Hess and Ostrom, 2007) . The second axis is the governance function performed by an institution over a layer, and here we adopt the eight design principles that Ostrom (1990) identified as characteristic of long-enduring self-governed institutions: the definition of clear boundaries around both resource and community, the congruence of appropriation and provision rules with local conditions, collective-choice arrangements that let those affected modify the rules, monitoring, graduated sanctions, accessible conflict-resolution mechanisms, minimal recognition of the right to organize, and, for larger systems, organization in nested enterprises. The Cartesian product of the two axes yields a classificatory space in which any commons-governed AI institution can be located, characterized by the layers it pools and the functions it performs over each. We populate this space by examining the published evidence layer by layer, by identifying ten recurrent institutional archetypes, and by reading their positions through a maturity matrix and a compar...

The contribution of the article is fourfold. It establishes commons-governed AI as a named institutional category and grounds it in the Ostromian tradition, giving the institutional critique of principlism a constructive answer. It proposes a taxonomy whose axes are inherited from a validated institutional grammar rather than constructed ad hoc, and it applies the taxonomy across the entire AI stack rather than to data alone. It treats the energy and carbon cost of computation as a commons-governance problem on a par with the governance of data and weights, rather than as an externality to be priced, drawing the measurement literature (Strubell et al. , 2019; Patterson et al. , 2021; Luccioni et al. , 2023) and the community-energy tradition (Wade et al. , 2025) into a single frame. And it sets out the failure modes that constrain the project, including the openwashing identified by Widder et al. (2024) , the compute bottleneck analyzed by Sastry et al. (2024) , and the classical risk of free-riding, together with a research and policy agenda for a polycentric AI commons.

The remainder of the article proceeds as follows. The next section reconstructs the commons tradition from Hardin and Ostrom through the knowledge-commons and commons-based peer-production literatures, isolating the conceptual moves the taxonomy depends on. The third section argues that the AI stack is a layered commons and defines the five resource layers. The fourth section constructs the taxonomy itself, stating its axes, its descriptive secondary dimensions, and the method by which it was built. The five sections that follow examine the data, compute, model, knowledge, and energy layers in turn. A synthesis section then assembles the institutional archetypes, presents the maturity matrix and the principle-by-principle comparison, and reads the durable combinations off the evidence. A penultimate pair of sections treats sectoral applications and the tensions and failure modes of the project, and a final section presents the research and policy agenda before concluding.

2 The Commons Tradition and Its Extension to Knowledge

The conceptual backbone of the taxonomy is the theory of the commons and its successive extensions from natural resources to knowledge, software, and data. We reconstruct that tradition here only to the depth the taxonomy requires, isolating the four moves on which the rest of the article depends: the distinction between open access and a governed commons, the design principles that characterize durable self-governance, the reframing of knowledge as a commons, and the account of commons-based peer production that explains how many AI commons are produced rather than merely held.

The first move is the separation of open access from the commons proper. The intuition that shared resources collapse under self-interest, formalized by Hardin (1968) , models a pasture to which access is unrestricted and in which no rule constrains appropriation. Under those conditions the marginal benefit of an additional animal accrues to the individual herder while the marginal cost of degradation is socialized, and the dominant strategy is to overgraze. The error that the commons tradition exists to correct is the identification of this open-access regime with all forms of shared ownership. Ostrom (1990) showed that the resources Hardin had in mind are in practice rarely open access; they are common-pool resources governed by bounded communities through self-determined institutions, and the empirical record of irrigation districts, inshore fisheries, alpine pastures, and forest commons contains many that have endured for centuries. A common-pool resource in this technical sense is one from which it is costly to exclude users but whose units are subtractable, so that one appropriator’s use diminishes what remains for others; the governance problem is to align appropriation and provision so that the resource is neither overused nor under-maintained. The distinction is load-bearing for the present article because it licenses treating the AI stack as a governance design problem rather than as a tragedy to be averted only through enclosure or top-down regulation.

The second move is analytical. From her comparative field studies Ostrom (1990) distilled a set of design principles that long-enduring self-governed common-pool-resource institutions tend to share, and that institutions which fail tend to violate. These principles are the columns of our taxonomy, so we state them in the form in which the taxonomy uses them.

Design Principle 1 (Clearly defined boundaries) .

The boundaries of the resource and the set of individuals or groups entitled to appropriate from it are clearly defined. For an AI commons this concerns both what counts as the pooled resource, a corpus, a weight set, a compute allocation, and who is a member of the governing community.

Design Principle 2 (Congruence between rules and local conditions) .

Appropriation rules restricting time, place, technology, and quantity are congruent with local conditions and with provision rules requiring contributions of labor, money, or materials. A commons that lets members take without contributing, or that imposes provision burdens unrelated to benefit, is unstable.

Design Principle 3 (Collective-choice arrangements) .

Most individuals affected by the operational rules can participate in modifying them. This is the principle that most sharply separates a commons from both a firm, where rules are set by ownership, and a regulator, where they are set by statute.

Design Principle 4 (Monitoring) .

Monitors who audit resource conditions and appropriator behavior are accountable to the appropriators or are the appropriators themselves. In the AI setting, monitoring maps onto transparency and auditability of data provenance, model behavior, and compute and energy use.

Design Principle 5 (Graduated sanctions) .

Appropriators who violate operational rules are subject to graduated sanctions, proportionate to the seriousness and context of the offense, administered by other appropriators or by accountable officials.

Design Principle 6 (Conflict-resolution mechanisms) .

Appropriators and their officials have rapid access to low-cost local arenas for resolving conflicts among themselves or with officials.

Design Principle 7 (Minimal recognition of rights to organize) .

The right of appropriators to devise their own institutions is not challenged by external governmental authorities. For AI commons this concerns the legal recognition of trusts, cooperatives, foundations, and licenses as vehicles of self-governance.

Design Principle 8 (Nested enterprises) .

For commons that are parts of larger systems, appropriation, provision, monitoring, enforcement, conflict resolution, and governance are organized in multiple nested layers. This is the principle of polycentricity, and it is the one that scales a local AI commons into an ecosystem.

The analytic engine beneath these principles is the Institutional Analysis and Development framework, which decomposes any governance situation into an action arena conditioned by the biophysical attributes of the resource, the attributes of the community, and the rules-in-use, and which resolves the rules into position, boundary, choice, aggregation, information, payoff, and scope rules (Ostrom, 2005, 2011) . The framework is a diagnostic scaffold rather than a predictive theory, and it supplies the dimensions along which our taxonomy compares institutions, so that the axes are inherited from a validated grammar rather than invented for the occasion.

The third move extends the commons from natural resources to knowledge. Hess and Ostrom (2007) argue that information and knowledge can be analyzed as commons even though their content is non-rival in consumption, because the infrastructures that store, curate, and transmit them, and the labor that maintains provenance and quality, are subtractable and congestible. This is the conceptual hinge that lets us treat a training corpus, a weight set, or a benchmark as a genuine commons rather than as a mere public good: copying a dataset does not diminish it, but curating, documenting, hosting, defending its licensing, and sustaining its provenance are rivalrous activities that demand provision rules. The Governing Knowledge Commons program of Madison et al. (2010) and Frischmann et al. (2014) sharpens this insight into a structured case-study method, observing that cultural and intellectual commons are constructed rather than naturally bounded and that their resources are frequently produced by the very sharing arrangement under study. Our per-archetype characterization mirrors this interrogation, asking of each commons-governed AI institution how its boundaries are constructed, who its contributor community is, what its openness rules are, and what outcomes it produces.

The fourth move concerns production. Benkler (2002) introduces commons-based peer production, arguing through the economics of the firm that radically decentralized, non-proprietary collaboration can outperform both firms and markets for information production when the work is modular, its components are of fine and heterogeneous granularity, and the cost of integrating contributions is low; free and open-source software is the canonical case, and the argument is developed into a general account of the networked information economy in Benkler (2006) . The empirical counterpart is the large-N study of free/libre and open-source projects by Schweik and English (2012) , which identifies the conditions under which open-source commons succeed rather than are abandoned. This strand matters because many AI commons are not merely governed pools but are produced by peer production, and the taxonomy must distinguish the governance of a resource from the peer-production process that creates it. Against this constructive literature stands the critical counter-narrative of enclosure. Boyle (2003) frames the expansion of intellectual property as a second enclosure movement fencing off the intangible public domain, and Lessig (2006) makes the foundational argument that technical architecture is itself a regulator, that code is law. Both are essenti...

A fifth move situates the commons tradition within the older intellectual lineage from which it descends, for the Ostromian programme is best read as the empirical formalization of a long current of non-statist and self-organizing political thought rather than as its origin. The proposition that durable order can arise from the free federation of autonomous units, without recourse to a sovereign centre, is the core of the mutualist and federalist arguments of Proudhon (1979) , and the nested, multi-layer scaling that Ostrom names polycentricity is the institutional descendant of that federative principle. The claim that cooperation, and not competition alone, is a primary factor in social organization was advanced by Kropotkin (1902) , and it supplies the motivational substrate that the design principles presuppose but do not themselves explain, namely why appropriators contribute provision when narrow self-interest would counsel free-riding. The confederal and ecological strand of Bookchin (1982) ties the self-government of resources to a critique of hierarchy and to ecological limits, prefiguring both the polycentric scaling of the eighth design principle and the treatment of energy and sustainability as governance problems internal to the commons rather than as external constraints. That such self-organization is not utopian but already pervasive in ordinary social life is the observation of Ward ([1973](https://arxiv.org/html...

3 A Genealogy of the Commons Idea, from Federation to Artificial Intelligence

The previous section read the commons tradition analytically, as the source of the design principles that form the columns of the taxonomy. It is worth reading the same tradition a second time chronologically, because the question that motivates this article, how a resource as novel as the AI stack could already have a mature vocabulary of collective governance waiting for it, is answered by the long history of the idea rather than by its logical structure. The vocabulary was not built for artificial intelligence; it was built over more than a century and a half for pastures, forests, software, and data, and it arrives at the AI stack as the latest term in a sequence. Figure 2 sets out that sequence in three periods, and we trace it here to establish that commons-governed artificial intelligence is the current chapter of a continuous argument and not a recent coinage.

The first period runs from the middle of the nineteenth century to the end of the twentieth and concerns natural resources and the political theory of self-organization. The thesis that durable order can arise from the free federation of autonomous units, without a sovereign centre, is stated as a principle of social organization by Proudhon (1979) , whose argument that legitimate order is built upward from federated units rather than imposed downward from a state is the distant ancestor of what Ostrom would later call polycentricity. The claim that cooperation is a factor of survival on a par with competition is advanced by Kropotkin (1902) , supplying the motivational premise, why members provide to a commons rather than free-ride, that the design principles presuppose but do not themselves prove. For most of the twentieth century these remained heterodox positions, and the dominant analytic frame was instead the one crystallized by Hardin (1968) , whose tragedy of the commons appeared to demonstrate that shared resources must be either privatized or policed. The decisive empirical correction is the comparative fieldwork of Ostrom (1990) , which established that the resources Hardin described are in practice governed by bounded communities under self-given rules, and which distilled the eight design principles; the complementary epistemic argument, that centrally legible schemes fa...

The second period, roughly the first decade and a half of the present century, is the migration of the commons from natural resources to information, and it is the period that makes the application to artificial intelligence possible at all. The economic argument that radically decentralized, non-proprietary collaboration can outproduce both firms and markets for information goods is made by Benkler (2002) and developed into a general account of the networked economy in Benkler (2006) . The counter-movement of enclosure is named in the same years: Boyle (2003) describes the expansion of intellectual property as a second enclosure of the intangible public domain, and Lessig (2004) documents how law and technology together lock down culture, an argument he sharpens into the thesis that architecture is itself a regulator, that code is law (Lessig, 2006) . The conceptual hinge on which this article turns is supplied in this period by Hess and Ostrom (2007) , who argue that knowledge can be analyzed as a commons despite the non-rivalry of its content, because the infrastructures and labor that curate and sustain it are subtractable. The empirical study of free and open-source software by Schweik and English (2012) and the case-study method of...

What the chronology shows that the analytic reconstruction cannot is the direction and the speed of transmission. The idea moved from pastures to bitstreams to model weights, each migration carrying the same core, a bounded community giving itself rules over a subtractable resource, into a domain its originators did not anticipate, and the migrations grew closer together as the digital substrate spread. The AI-era instances of the third period arrived, however, without the comparative grammar that the first two periods had developed, each named in the vocabulary of its own subfield and studied in isolation. The contribution of the present article is to supply that grammar, by recognizing the third-period instances as the continuation of the first two and by classifying them with the apparatus those periods produced. With the lineage of the idea established, we turn to the object it is now being asked to govern, the layered AI stack.

4 The AI Stack as a Layered Commons

To classify how AI is governed in common we must first fix what is being governed. The contemporary AI supply chain is conventionally decomposed into three core layers, compute, data, and models, where compute denotes the physical and software infrastructure of accelerators, interconnect, and data centers, data denotes the corpora on which models are trained and evaluated, and models denotes the trained artifacts and the weights that parameterize them. We retain this decomposition but extend it in two directions that the commons frame makes necessary. Downward, we add energy as the physical substrate on which compute runs, because at current scale the electricity and water drawn by AI training and inference have become shared-resource conflicts in their own right (International Energy Agency, 2025; Masanet et al. , 2020) . Laterally, we separate a knowledge and evaluation layer from the model layer, comprising the benchmarks, leaderboards, documentation standards, and open toolchains that are produced by peer production and that determine what counts as progress. The result is a five-layer view of the AI stack as a stratified commons, each layer of which we now define and argue is genuinely commons-like.

Definition 1 (Commons-governed AI institution) .

A commons-governed AI institution is an arrangement in which one or more resource layers of the AI stack are pooled and stewarded as a shared resource whose access, contribution, and benefit are determined by a defined community under self-given rules, rather than by unilateral private control or top-down statutory command. Equivalently, in the terms of Section 2, it is an institution whose decision rights over an AI resource are held collectively and whose operation can be characterized by Ostrom’s design principles.

The qualifier that does the work in this definition is subtractability. A resource need not be physically rival to be a commons; it need only be subtractable in some governance-relevant dimension, so that uncontrolled appropriation or insufficient provision degrades it (Hess and Ostrom, 2007) . We argue layer by layer that each of the five satisfies this criterion.

The data layer comprises the training and evaluation corpora. Their informational content is non-rival in copying, but the activities that make a corpus usable, curation, deduplication, documentation, consent management, license defense, and the maintenance of provenance, are rivalrous in labor and congestible in infrastructure, and the supply of high-quality openly licensed data is in fact shrinking under enclosure pressure as rights-holders restrict reuse for AI training (Purtova and van Maanen, 2024) . A corpus left ungoverned degrades through quality deterioration, contamination, and the erosion of consent, which is exactly the provision-and-appropriation problem Ostrom’s principles address.

The compute layer comprises accelerators, clusters, and the cloud infrastructure that schedules them. Compute is the most clearly rival of the five layers and, as Sastry et al. (2024) argue, the most governable node of the supply chain, because it is detectable, excludable, quantifiable, and produced through an extremely concentrated manufacturing pipeline. These same properties make compute both the most natural candidate for enclosure and the layer where a commons is hardest to sustain: a resource that is easy to meter and to deny is a resource around which exclusion is the default. The access asymmetry this produces, the compute divide between the few organizations that command frontier-scale clusters and everyone else, is the inequality that public-compute and compute-commons proposals exist to redress (Heim et al. , 2024; National Artificial Intelligence Research Resource Task Force, 2023) .

The model layer comprises trained weights and the interfaces through which they are served. Weights are non-rival once released, but their governance is subtractable through the architectural channel identified by Lessig (2006) : a developer chooses a release modality, from a closed inference API through gated weights to fully open weights, and that choice is a governance act that the community around the model may or may not control. The transparency of the upstream resources that produced a model, its data, its data labor, and its computational cost, is itself a scarce and contested good, as the low scores on exactly those indicators in the Foundation Model Transparency Index demonstrate (Bommasani et al. , 2023, 2024) .

The knowledge and evaluation layer comprises benchmarks, evaluation suites, documentation standards such as model cards (Mitchell et al. , 2019) , and the open software toolchains on which the field runs. This is the layer most fully explained by commons-based peer production (Benkler, 2002) : benchmarks and libraries are produced by loosely coordinated contributors, are non-rival in use, and are nonetheless subtractable in maintenance, because an unmaintained benchmark becomes contaminated and an unmaintained library bit-rots. The governance question is who sets the standard of progress and who bears the upkeep.

The energy layer comprises the electricity, and relatedly the water for cooling, consumed by training and inference. Energy is unambiguously rival and its appropriation is a classical commons problem of siting, grid impact, and carbon intensity. Per-model accounting has made the cost legible, from the training-time emissions measured by Strubell et al. (2019) and Patterson et al. (2021) to the full life-cycle assessment of Luccioni et al. (2023) , and the macro-scale demand projections of International Energy Agency (2025) turn that cost into a governance problem at the level of national grids. Treating energy as a layer of the AI commons rather than as an external cost is one of the article’s organizing choices, and it connects AI governance to the established community-energy commons tradition (Wade et al. , 2025) .

These five layers are not independent. Energy constrains compute, compute and data jointly produce models, and the knowledge layer determines how all of the others are measured and improved. A consequence we develop in Section 11 is that commons governance applied to a single layer in isolation is fragile: open weights served from enclosed compute trained on opaque data leave the underlying concentration of power intact, which is precisely the critique that Widder et al. (2024) level at nominal openness. The layered view therefore does double duty, fixing the first axis of the taxonomy and motivating its central empirical claim that durable AI commons are cross-layer rather than single-layer.

5 Constructing the Taxonomy

A taxonomy is useful to the degree that its dimensions are conceptually grounded, mutually exclusive enough to classify, and collectively exhaustive enough to cover the phenomenon. We build ours on two primary dimensions, both inherited from the commons tradition rather than constructed for the occasion, and we record for each institution a set of secondary descriptive attributes that capture variation the two primary axes do not. The construction follows the iterative logic of established taxonomy-development method (Nickerson et al. , 2013) , alternating between the conceptual derivation of dimensions from the Ostromian grammar and the empirical inspection of the institutions examined in Sections 610, and it adopts the interrogation structure of the Governing Knowledge Commons case-study method (Frischmann et al. , 2014) , which asks of every commons how its resource and community boundaries are constructed, what its rules-in-use are, and what outcomes it produces.

The first primary dimension is the resource layer of the AI stack that an institution pools, taking values in the five-element set fixed in Section 4: data, compute, models, knowledge and evaluation, and energy. A given institution may pool more than one layer, and indeed Section 11 argues that the durable institutions are precisely those that pool several; the layer dimension is therefore polythetic, recording the subset of layers an institution governs rather than assigning it to a single class.

The second primary dimension is the governance function an institution performs, taking values in the eight design principles of Section 2: boundary definition, congruence of appropriation and provision, collective choice, monitoring, graduated sanctions, conflict resolution, recognition of the right to organize, and nested polycentric scaling. For each (layer, function) pair an institution either does or does not implement the function over the layer, and the strength with which it does so is what the maturity matrix of Section 11 records. Crossing the five layers with the eight functions yields the forty-cell classificatory space of Figure 4, in which any commons-governed AI institution occupies a region rather than a point: a federated health consortium, for example, is concentrated in the data and compute rows and is strong on boundaries, monitoring, and collective choice but weak on graduated sanctions, while a public compute initiative occupies the compute row and is strong on monitoring and recognition but weak on collective choice.

Two primary axes do not exhaust the variation among commons-governed AI institutions, so we record four secondary descriptive attributes that the layer sections fill in and that the synthesis collects in Table 2. The first is the position on the openness spectrum, which ranges from a closed proprietary resource, through a club good available to members on terms, through genuinely open access, to a fully commons-governed resource that is both open and collectively stewarded; the distinction between mere open access and a governed commons, drawn in Section 2, is exactly the distinction the upper end of this spectrum encodes, and Section 8 shows that openness without collective governance is the failure mode Widder et al. (2024) diagnose. The second is the legal vehicle through which the institution holds rights and bears duties, ranging across the trust, the cooperative, the foundation or non-profit association, the public agency or intergovernmental body, the contractual license, and the decentralized autonomous organization (Hassan and De Filippi, 2021) ; the legal vehicle is the concrete form Ostrom’s seventh principle, recognition of the right to organize, takes in practice. The third is the rule for benefit distribution, which determines whether value created by the commons flows to contributors, to all members equally, to the broader public...

The classification produced by these dimensions is polythetic and morphological rather than monothetic: an institution is characterized by the profile of values it takes across layers, functions, and secondary attributes, not assigned to one of a small number of mutually exclusive boxes. This is a deliberate choice forced by the phenomenon, because the institutions examined below recombine the same governance functions over different layers, and a monothetic scheme would either multiply classes uncontrollably or suppress the cross-layer combination that is the empirical heart of the subject. The ten archetypes named in Section 11 are not the classes of the taxonomy but recurrent regions of the space, useful as landmarks precisely because real institutions cluster near them while varying in the secondary attributes.

As a reading map for the evidence that follows, Table 1 groups the literature the taxonomy draws on by the region of the classificatory space each body of work informs, separating the conceptual foundations and the AI-governance context that supply the two axes from the five resource layers that the next sections examine in turn. The grouping is itself an application of the taxonomy, in that a work is placed by the layer and function it speaks to rather than by its disciplinary origin, and the recurrence of several works across cells is the bibliographic trace of the cross-layer character that Section 11 argues is constitutive of the durable commons.

With the axes fixed, we turn to the evidence, examining each resource layer in the order data, compute, models, knowledge, and energy.

6 The Data Layer

The data layer is the most institutionally developed of the five, because the governance of data predates the foundation-model era and arrives in the AI commons with a mature vocabulary of trusts, cooperatives, and sovereignty regimes. The organizing question of the layer is who holds the rights to pool, curate, and license a corpus, and to whose benefit the resulting models flow. Four families of arrangement answer it in commons terms, and each instantiates a recognizable subset of Ostrom’s design principles over the data row of the taxonomy.

The data trust is a fiduciary arrangement in which data subjects pool rights of control over their data and vest them in a trustee who is legally bound to exercise those rights in the subjects’ interest. Delacroix and Lawrence (2019) develop the bottom-up data trust as a deliberate alternative to a one-size-fits-all default, arguing for an ecology of trusts with differing terms among which subjects can choose and to which they can defect, so that the trust structure supplies the collective-choice and conflict-resolution functions that bilateral consent between an individual and a platform cannot. In taxonomic terms the trust is strong on recognition of the right to organize, because the trust is a well-understood legal vehicle, and on boundary definition, because membership and the pooled resource are defined by the trust instrument, while its collective-choice strength depends on whether beneficiaries can instruct the trustee or merely exit. The trust answers the fiduciary-duty objection that Mittelstadt (2019) raises against principlism, supplying exactly the enforceable duty that a list of principles lacks.

The data cooperative places curation and analytics under member ownership. Hardjono and Pentland (2019) model it on the credit union, a member-owned fiduciary that aggregates members’ personal data, runs analytics on their behalf, and returns the resulting value to the membership, so that the cooperative is strong on collective choice and on benefit distribution to contributors, the two functions that most sharply distinguish it from a firm. The cooperative form realizes Ostrom’s third principle directly, because members participate in modifying the operational rules through democratic governance, and it realizes the second principle, congruence of appropriation and provision, because members both contribute data and share in the surplus. Its characteristic weakness is on graduated sanctions and monitoring at scale, since a large cooperative struggles to audit member behavior, and on the provision of the compute needed to act on the pooled data, which is why Section 11 reads the cooperative as a data-row institution that becomes durable only when coupled to a compute-row arrangement.

The closest prior taxonomy of the layer is that of Micheli et al. (2020) , who identify four emerging models of data governance, data-sharing pools, data cooperatives, public data trusts, and personal data sovereignty, classified by who holds power and to whose benefit data flows. That taxonomy is the proximate ancestor of the present one and the reason we position carefully against it: it is organized by actor and power rather than by commons-governance mechanism, it predates the foundation-model era, and it treats data in general rather than AI training assets in particular. Our contribution at this layer is to re-read its four models as profiles over the eight Ostromian functions and to extend the analysis from personal data to the training corpora and provenance infrastructures that foundation models consume. The survey of Purtova and van Maanen (2024) supplies the bridge, distinguishing data as an economic good from data as a commons and assessing the governance models against the common-pool-resource tradition, and it documents the enclosure pressure, the shrinking supply of openly licensed data, that makes a data commons urgent rather than optional.

The cultural and community strand is where the data layer most clearly exceeds individual consent. The CARE principles for Indigenous data governance, collective benefit, authority to control, responsibility, and ethics, articulated by Carroll et al. (2020) , assert that data about a people are governed by that people as a collective, not merely by the individuals it describes, and they were formulated as a complement to the stewardship-oriented FAIR principles of Wilkinson et al. (2016) , which make data findable, accessible, interoperable, and reusable but say nothing about who decides. Indigenous data sovereignty (Kukutai and Taylor, 2016) is the broader movement of which CARE is the principled expression, and it grounds the taxonomy’s claim that commons governance is not reducible to open access: a corpus can be fully open and still violate collective authority, and a commons can legitimately restrict appropriation to honor it. This strand is strong on collective choice and on boundary definition in the communal sense, and it supplies the normative content that distinguishes a commons from a public good.

The empirical referents of the layer are the large open corpora on which the claim that AI training data can be commons-governed must ultimately be tested. The Common Voice project of Ardila et al. (2020) assembles a massively multilingual speech corpus through voluntary contribution and a permissive dedication to the public domain, and it is among the cleanest existing instances of commons-based peer production applied to AI training data, strong on provision through contribution and on open boundaries. The LAION-5B image-text dataset of Schuhmann et al. (2022) , by contrast, illustrates the governance questions a large open corpus raises rather than settles, since its openness coexists with contested questions of consent, provenance, and harmful content that a genuine data commons would have to govern through monitoring and graduated sanctions it did not originally possess. Read together, the two cases mark the range of the data layer: openness is necessary for a data commons but not sufficient, and the difference between Common Voice and the controversies around large scraped corpora is precisely the difference between a governed commons and mere open access that Section 2 insists upon. The practitioner literature, including the GovLab blueprint for unlocking new data commons for AI (Chafetz et al. , 2024) and the Open Future proposal for commons-based dataset governance (Open Future F...

7 The Compute Layer

Compute is the layer at which the gap between the promise of an AI commons and the reality of concentration is widest. Sastry et al. (2024) argue that computing power is a uniquely governable node of the AI supply chain because it is detectable, excludable, quantifiable, and produced through a manufacturing pipeline of extreme concentration, in which a handful of firms control design, fabrication, and the critical equipment upstream of them. The argument cuts two ways for the commons project. The properties that make compute a lever for state and corporate control, excludability and concentration, are the same properties that make a compute commons both necessary and difficult, because a resource that is trivial to meter and to deny is a resource around which enclosure is the default equilibrium and pooling is the exception that must be deliberately constructed and defended. Heim et al. (2024) sharpen the institutional picture by locating cloud compute providers as regulatory intermediaries, with latent capacities as securers, record keepers, verifiers, and enforcers, and we read those capacities as the plumbing through which either enclosure or commons-style stewardship must operate: the same provider that can enforce a state export control can, under different governance, supply the monitoring and graduated sanctions a compute commons requires.

The asymmetry these analyses describe is the compute divide, the concentration of frontier-scale training capacity in a few organizations while public-interest users, academic researchers, smaller firms, and the global majority operate at orders of magnitude less. The divide is not merely a matter of cost but of governance, because the organizations that command the compute also command the agenda of what is built and evaluated, and it is the inequality that the public-compute proposals of the layer exist to redress. The United States National AI Research Resource is the most developed national articulation (National Artificial Intelligence Research Resource Task Force, 2023) , proposing a cooperative-stewardship model in which a federal mix of computational and data resources is made accessible through an integrated portal with the explicit aims of spurring innovation, diversifying the talent base, improving capacity, and advancing trustworthy AI. In taxonomic terms the Resource is a state-anchored quasi-commons: it pools provision and exercises collective choice through an advisory governance structure, and it is strong on monitoring and on recognition of the right to organize, but it stops short of the user self-governance that Ostrom’s third principle places at the center, since eligibility and allocation are set administratively rather than by the user community. That gap is exactly the distinction the taxonomy uses to separate a public-provision cell from a fully commons-governed one, and it is a productive gap rather t...

The European counterpart supplies a second institutional form. The EuroHPC Joint Undertaking, which pools publicly owned supercomputing capacity across member states, was extended by a 2024 amendment to establish AI Factories that repurpose that capacity for AI training and offer it to European users (European High Performance Computing Joint Undertaking, 2024) . The Undertaking is federated across national and Union levels, which makes it the clearest existing instance of Ostrom’s eighth principle, nested enterprise, applied to compute: appropriation and provision are organized in layers, with national centers nested within a Union-level coordinating body. Together the United States and European initiatives establish that large-scale public compute provision is no longer hypothetical, which lets the taxonomy treat the compute commons as an empirical category rather than a thought experiment, while their shared distance from user self-governance marks the frontier the agenda of Section 14 addresses.

A distinct and technically grounded route to a compute commons does not pool hardware at all but pools learning. Federated learning, introduced by McMahan et al. (2017) as communication-efficient training of a shared model over decentralized data through iterative averaging of locally computed updates, and surveyed comprehensively by Kairouz et al. (2021) , lets a community train a common model without centralizing either the raw data or, in principle, the compute, since each participant contributes the computation local to its own data. Federated learning is therefore the substrate that makes a commons of learning feasible where a commons of raw data is blocked by privacy, regulation, or sovereignty, and it recurs in the sectoral applications of Section 12, most prominently in health, where data cannot leave the institution that holds it. In taxonomic terms a federated consortium is strong on boundary definition and on monitoring, because participation and contribution are explicit and auditable, and its collective-choice strength depends on whether the federation’s governance lets participants set the training objective and the aggregation rule or merely opt in to one set by a coordinator. Its characteristic weaknesses are graduated sanctions, since excluding a misbehaving participant is the only available sanction and is rarely graduated, and the governance of the aggregate model’s weights, which returns us to the model layer. The...

8 The Model Layer

The model layer governs trained weights and the interfaces through which they are served, and it is the layer on which the meaning of openness is most contested. The foundation-model paradigm named by Bommasani et al. (2021) concentrated capability into a small number of large artifacts whose governance is determined by a single developer choice, the release modality, which ranges along a spectrum from a fully closed inference interface to fully open weights and beyond. Because that choice is an exercise of the architectural power that Lessig (2006) identified as a form of law, the model layer is where the taxonomy’s secondary openness dimension does the most work, and where the distinction between mere open access and a governed commons becomes a practical design problem rather than a conceptual one.

Figure 5 arranges the layer along that spectrum. At the closed end a model is served through an API that exposes behavior but not weights, and governance is wholly internal to the developer. A gated release discloses weights to approved parties under contractual terms, a club good in the sense of Section 5. An open-weights release publishes the parameters for anyone to run and adapt, which is genuine open access but, by itself, says nothing about who governs the data, the compute, or the documentation that produced them. The Open Source Initiative’s Open Source AI Definition raises the bar by requiring that the freedoms to use, study, modify, and share extend to the components needed to exercise them, including sufficient information about training data (Open Source Initiative, 2024) , and only at the far end does a model become commons-governed in the full sense, openly licensed and also collectively stewarded across the layers that produced it.

The instruments that would let a community govern this layer rather than merely consume its outputs are transparency and documentation. The Foundation Model Transparency Index of Bommasani et al. (2023) and its 2024 successor Bommasani et al. (2024) score developers across upstream resources including data, data labor, and compute, and the empirically telling result is that developers score worst precisely on those upstream indicators, the commons-relevant layers of the stack, even when they score well on downstream access. Model cards (Mitchell et al. , 2019) supply the documentation primitive that turns a released model into an accountable one by recording intended use, training data, evaluation, and known limitations, and they are the model-layer analogue of the monitoring principle. The collaborative exemplar of the layer is the BigScience workshop’s BLOOM, a 176-billion-parameter multilingual model built by an open scientific collaboration on a publicly granted national supercomputer and released under a Responsible AI License (BigScience Workshop et al. , 2022) . BLOOM instantiates a compute-commons-plus-open-weights pathway, and paired with the life-cycle carbon assessment of the same artifact by Luccioni et al. (2023) , it demonstrates that a single model can be collectively built, openly licensed, and transparently measured...

The critical counterweight, and the reason the openness spectrum is a secondary axis rather than the primary one, is the argument of Widder et al. (2024) that nominally open AI systems are frequently closed in substance and that openness alone does not reduce the concentration of power. A model released with open weights but trained on undisclosed data using compute available only to its developer has redistributed an artifact without redistributing the capacity to produce or contest it. The position paper of Kapoor et al. (2024) on the societal impact of open foundation models gives the most careful benefit-and-risk accounting and argues that the marginal-risk evidence for the most cited misuse vectors of open weights remains thin, which the taxonomy uses to resist both a simple equation of openness with safety and its opposite. The combined lesson of the model layer is the article’s organizing constraint: a model commons is not constituted by open weights, but by open weights conjoined with collective governance of the data that trained them, the compute that produced them, the documentation that describes them, and the energy that powered them. Openness is the entry condition; the commons is the conjunction.

9 The Knowledge and Evaluation Layer

The knowle...