Think Inside the Box for Software Libraries
A White Paper on Co-Packaging Code and AI-Optimized Documentation as an Industry Standard
Authors: AI Documentation Working Group - Andres IGEA OVIEDO (andres.igeaoviedo@amadeus.com), Andrea SCORTI (andrea.scorti@amadeus.com)
Date: 29th of September 2026
Version: 1.1
Abstract
The rapid adoption of AI coding assistants — GitHub Copilot, Claude, Codex, and others — has fundamentally changed how developers consume software libraries. Yet the ecosystem’s documentation infrastructure remains anchored in a pre-AI paradigm: code ships in a package, documentation lives elsewhere. This disconnect forces AI agents to operate without the contextual knowledge they need, producing incorrect API usage, outdated patterns, and hallucinated interfaces.
This paper argues for a new industry standard: co-packaged documentation (CoDoc) — the practice of shipping structured, AI-optimized documentation alongside library code within the same distributable artifact. We call this the Think Inside the Box Principle: everything the consumer needs arrives in one box.
We present a reference implementation from the Design Factory design system, demonstrate that the approach works across package ecosystems (npm, pip, Maven), and propose a minimal specification that library authors can adopt today.
We also address critical concerns around security (prompt injection via co-packaged content), forward-compatibility with evolving agent architectures, scalability across deep dependency graphs, and the practical challenges of documentation generation pipelines.
1. Introduction: The AI-Driven Development Shift
Software development is undergoing its most significant workflow transformation since the introduction of the integrated development environment. AI coding assistants are no longer novelties — they are daily tools for millions of developers. GitHub reports that Copilot generates a significant share of new code in enabled repositories. Anthropic’s Claude, OpenAI’s Codex, and a growing ecosystem of AI agents are being embedded into every stage of the development lifecycle.
These agents share a common trait: they are only as effective as the context they can access. An AI assistant writing React code with full access to the React documentation produces dramatically better output than one relying solely on its training data. Training data goes stale; documentation is versioned and current.
Yet today, when a developer installs a library — whether via npm install, pip install, or a Maven dependency — the AI assistant working in that project receives the library’s code but almost never its documentation. The documentation exists, but it lives on a website, behind a search engine, in a format optimized for human browsing rather than machine consumption.
This is the gap we propose to close.
2. The Current State: A Separated World
2.1 How Libraries Ship Today
Across every major package ecosystem, the convention is the same: the package contains code; the documentation is hosted elsewhere.
| Ecosystem | Package Contents | Documentation Location |
|---|---|---|
| npm | JavaScript/TypeScript source, package.json, README |
Dedicated docs site, GitHub wiki, or MDN |
| PyPI | Python source, setup.py/pyproject.toml, README |
Read the Docs, Sphinx-hosted site, or GitHub pages |
| Maven | Compiled JARs, POM metadata | Javadoc on Maven Central, project wiki, or vendor site |
| NuGet | .NET assemblies, XML doc comments | Microsoft Learn, GitHub, or custom docs portals |
| Crates.io | Rust source, Cargo.toml |
docs.rs (auto-generated) |
The README file — the one piece of documentation that does travel with the package — is typically a brief overview: a logo, a one-liner description, installation instructions, and a link to “full documentation” hosted online.
This separation made sense in a human-centric world. Developers browse documentation in a web browser with search, navigation, and rich formatting. Shipping a full documentation site inside every node_modules dependency would have been wasteful.
But the consumer has changed. The primary documentation consumer is increasingly an AI agent running locally in the developer’s IDE, and the agent’s capabilities in browsing websites is not as robust as the capabilities it has for searching, navigating and reading files in a local workspace.
2.2 What AI Agents Can and Cannot Do
Modern AI coding assistants operate within a file-based context window. They can:
- Read files from the local filesystem (source code, configuration, markdown)
- Navigate directory structures
- Follow references between files
- Process structured and unstructured text
They typically have difficulty with, cannot or should not need to:
- Browse the internet during code generation
- Authenticate with documentation portals
- Parse JavaScript-rendered single-page application doc sites
- Maintain persistent knowledge across sessions about external APIs
When an AI agent encounters an unfamiliar library in a project’s dependencies, it has three options:
- Rely on training data — which may be months or years out of date, missing recent API changes, deprecations, or new features.
- Search the web — which adds latency, requires network access (often unavailable in corporate environments), and returns HTML pages optimized for human reading.
- Read local files — fast, always available, version-matched to the installed library, and optimized for machine consumption.
Option 3 is clearly superior. But today, there is almost nothing to read.
2.3 The Consequences of the Gap
The documentation gap produces observable failures in AI-assisted development:
- Hallucinated APIs: The agent invents function signatures, props, or class methods that don’t exist in the installed version, because it is interpolating from stale training data.
- Deprecated patterns: The agent uses patterns from an older version of the library because its training data predates the current release.
- Missing context: The agent cannot recommend the idiomatic way to use a component because it has no access to usage guidelines, do/don’t rules, or design system constraints.
- Repeated correction cycles: The developer must manually correct the agent’s output, defeating the productivity gain that AI assistance promises.
These are not edge cases. They are the daily experience of developers using AI tools with any library that has evolved since the agent’s training cutoff. Preliminary observations from Design Factory’s internal adoption suggest that co-packaged documentation significantly reduces these failures, though rigorous benchmarking across diverse libraries remains an open research direction.
To close this evidence gap, the authors intend to develop CoDocBench, a public evaluation suite measuring AI-agent performance on library-usage tasks (for example, API correctness and deprecated-pattern rates) with and without co-packaged documentation, across multiple libraries and package ecosystems. CoDocBench is intended to provide measurable success criteria for the standard proposed in this paper and to invite external validation of its claims.
3. Related Work and Adjacent Conventions
CoDoc does not exist in a vacuum. Several existing conventions already address fragments of the problem this paper targets. This section surveys the closest adjacent work. Analysing this work reveals an emerging pattern: existing efforts focus on the format problem (how to write documentation for AI consumption) or on providing hosted retrieval (serving documentation from a registry or third party), but none solve the delivery and logistics problem: getting version-matched documentation onto the agent’s filesystem, inside the package artifact.
3.1 llms.txt: The Right Format in the Wrong Location
The closest existing convention is llms.txt, a proposal that websites publish a markdown file /llms.txt at the site’s root or at any path within it, with an optional /llms-full.txt containing the entire collection of the site’s core content so that AI agents can consume its content without parsing HTML. It has seen adoption across some documentation sites and developer tools.
llms.txt gets the format right: curated markdown, an index enabling selective reading, prose that can be mixed with code examples. These are some of the same properties Section 8.4 specifies for CoDoc documentation files, and the convergence is encouraging. The ecosystem is independently arriving at markdown as the canonical AI-consumable format.
Where llms.txt differs is location, and location is precisely the problem. An llms.txt file lives on a website, so the versioning, offline, and determinism arguments of Section 10.6 apply to it directly:
- Version mismatch. A site’s
llms-full.txtdescribes the version the site currently hosts, not necessarily the version installed in the developer’s project. - Availability. It requires network access and fails in air-gapped environments, CI pipelines, and restricted corporate networks.
- Determinism and trust. Its content can change between requests, and it is untrusted, unaudited web content the agent must ingest.
The two conventions are complementary: llms.txt answers “how should a website serve AI agents?”, while CoDoc answers “how should a package serve AI agents?”.
3.2 Hosted Documentation Intermediaries
A growing category of services such as Context7, DeepWiki, and similar documentation-serving layers index library documentation or source repositories and serve the result to AI agents, typically through MCP tools or web interfaces. These services demonstrate real demand for agent-consumable documentation and are genuinely useful stopgaps.
In the framing of this paper, however, they are probabilistic middleware between the agent and the knowledge it requires (Section 6.2):
- Retrieval is not guaranteed. The agent must decide to query the service, the service must have indexed the library, and the indexed version must match the installed one. Each step is an independent failure point.
- Third-party maintenance repeats history. The content is curated by someone other than the library author, on infrastructure the author does not control. This is the DefinitelyTyped logistics problem again (Section 12): decoupled maintenance drifts silently.
- The trust boundary widens. Every hosted intermediary is another party whose content the agent ingests and whose availability the workflow depends on.
These services are a valuable bridge for the long tail of libraries that ship no agent-consumable documentation at all. But a bridge is not a foundation. The authoritative, version-matched source should ship from the library author, inside the package.
3.3 Agent Instruction Conventions: Complementary Infrastructure
A parallel set of conventions has emerged for instructing agents how to behave inside a project: AGENTS.md, .github/copilot-instructions.md, Cursor rules, and their tool-specific variants. CoDoc does not compete with these conventions; it gives them something stable to point at.
Instruction files are pointers; co-packaged documentation is the pointed-at content. Without CoDoc, an instruction file that wants to ground the agent in a library’s API can only point at a URL — inheriting the version-mismatch, availability, and determinism problems of Section 10.6. With CoDoc, the same instruction file can point out to the agent that all software libraries have co-packaged documentation alongside the library code within the same distributable artifact. CoDoc scales elegantly and avoids the hardcoding of URLs anywhere in a software engineering project.
This is exactly the two-tier discovery model defined in Section 8.5: auto-discovery of docs/ directories as the primary mechanism, instruction-file integration as the supplementary channel.
3.4 Summary
| Convention | AI-optimized format | Version-matched | Offline-capable | Maintained at the source |
|---|---|---|---|---|
llms.txt |
Yes | Not guaranteed | No | Yes |
| Hosted intermediaries (Context7, DeepWiki) | Yes | Sometimes (indexed version) | No | No (third party) |
Instruction files (AGENTS.md, etc.) |
N/A — instructions, not documentation | N/A | Yes | N/A (consumer-authored) |
| CoDoc | Yes | Yes | Yes | Yes |
No existing convention delivers version-matched, AI-optimized documentation inside the package artifact. This specific combination is CoDoc’s contribution.
4. The Think Inside the Box Principle
Some physical goods companies’ success is built on a deceptively simple idea: everything you need comes in one box. For instance, when you purchase a bookcase, the box contains the panels, the shelves, the screws, the dowels, the cam locks, and — critically — the assembly instructions. You do not need to visit a website to look up how to assemble it. You do not need to search YouTube for a tutorial. The instructions are right there, designed to be consumed alongside the product, version-matched and complete.
This principle has a direct analog in software:
Think Inside the Box for Software Libraries: A package should contain everything an agent needs to correctly use the library — the code and its documentation — in the same distributable artifact.
The analogy is precise:
| Physical goods company | Software Library |
|---|---|
| Product (panels, shelves) | Source code, compiled artifacts |
| Assembly and usage instructions | API docs, usage guides, examples |
| The box | The package (npm tarball, wheel, JAR) |
| The customer | The AI coding agent |
| The store | The package registry (npm, PyPI, Maven Central) |
Just as instructions from newly purchased furniture are designed for the person assembling it and using it — not the designer who created it — co-packaged documentation should be designed for the agent consuming the library, not the author who wrote it.
The key insight is that the consumer of library documentation is changing. Historically, documentation was written for human developers who read it in a browser. Increasingly, documentation must also serve AI agents that read it as local files. The Think Inside the Box Principle asks library authors to serve this new consumer by including documentation in the delivery, not just linking to it.
5. Design Factory: A Reference Implementation
The Design Factory design system provides a concrete implementation of co-packaged documentation. Analyzing its architecture reveals patterns that generalize across ecosystems.
5.1 What Ships in the Package
Design Factory publishes the following structure inside its npm package:
@design-factory/design-factory/
├── .ai/
│ ├── index.md # Component catalog (56 components)
│ ├── design-tokens.md # 4-tier token reference
│ ├── foundations.md # Typography, spacing, grid
│ └── docs/
│ ├── components/{slug}/
│ │ ├── api.md # Selectors, inputs, outputs
│ │ ├── overview.md # Component overview
│ │ ├── examples.md # Usage examples with do/don't
│ │ ├── guidelines.md # Usage guidelines
│ │ └── accessibility.md # Accessibility guidance
│ └── demos/{slug}/{variant}/
│ ├── component-variant.ts # Working demo source
│ └── component-variant.html # Demo template
This is the documentation, versioned and shipped with the code. When a developer runs npm install @design-factory/design-factory, they receive not just the Angular components but a complete, structured knowledge base that any AI agent can navigate.
A note on the
.ai/directory name. Design Factory’s first implementation predates the current CoDoc standard proposed in Section 8, which specifies a plaindocs/directory instead of a hidden.ai/directory. Early feedback on the.ai/convention recommended a location that is visible and accessible to both humans and AI agents: hidden dot-directories are easy to overlook in file explorers, and some tools exclude them from search and indexing by default. Design Factory will update its package structure to conform to the standard proposed in this paper, and is in the process of becoming fully open-source, which will make its documentation generation pipeline available as a public reference implementation.
5.2 How the Agent Uses It
The integration works through two complementary discovery mechanisms:
Auto-discovery (preferred). AI tool vendors are encouraged to automatically detect .ai/ directories in installed dependencies (see Section 11.4). When an agent encounters an unfamiliar API, it scans for a .ai/index.md in the relevant package and navigates from there. This requires zero configuration from the consuming developer.
Explicit instructions (supplementary). For tools that do not yet support auto-discovery, a standard instructions file (AGENTS.md) in the project root can direct the agent to .ai/ folders in specific dependencies:
- The AI agent reads
AGENTS.mdin the project root. - The rules direct it to the
.ai/folder in the installed package. - The agent reads
index.mdto understand what components exist. - For a specific task (e.g., building a form with a datepicker), the agent selectively reads only the relevant docs — API, examples, guidelines.
- The agent produces code using the correct components, tokens, and patterns.
No web browsing. No external tools. No MCP servers. No infrastructure. Just files.
5.3 Key Design Decisions That Generalize
Several decisions in the Design Factory implementation reflect principles that apply to any library:
Plain markdown over structured formats. LLMs are trained on vast amounts of markdown. It’s the format they understand best, allows mixing prose with code examples, and is readable by both humans and machines. JSON schemas or YAML configs would require parsing layers with no added benefit.
Dynamic discovery over front-loaded injection. Rather than injecting all documentation into the AI prompt, the system provides an index file. The agent reads the index, identifies what’s relevant, and selectively reads only the docs it needs. This respects context window limits and mirrors how a competent developer navigates documentation.
Single source of truth. The shipped documentation is generated from the same sources that power the human-facing documentation portal. There is no separate “AI version” to maintain. This eliminates documentation drift — when the library updates, the docs update with it.
6. Why Static Files Are the Pragmatic Starting Point
A natural objection is: “Wouldn’t it be better to expose documentation through an MCP server, a dedicated tool, or a documentation API?” We argue that static files should be the foundation, and the reasoning is pragmatic.
6.1 Against MCP Servers for Documentation
The Model Context Protocol (MCP) allows AI agents to call external tools during a conversation. An MCP server for a library could offer tools like get_component_docs(name) or search_api(query). However we have to highlight the following disadvantages:
- Context pollution: Every MCP server injects its tool schemas into the system prompt, consuming context window space on every conversation — even those unrelated to the library.
- Round-trip overhead: Each MCP tool call requires the agent to decide to call, format the request, wait for the response, and parse the result. File reads are a primitive every agent already has.
- Operational burden: An MCP server requires installation, configuration, and a running process. Markdown files in the package require nothing.
- MCP is for operations, not reference. MCP excels when the agent needs to do things — query a database, manipulate a design tool. Reading documentation is retrieval, not action. The agent already knows how to read files.
6.2 Against Skills and Prompt Templates
Skills (pre-written prompt expansions) inject multi-step workflows when triggered. A “use library X” skill could instruct the agent to read docs in a specific order. However:
- Too rigid for contextual work. Library usage is deeply contextual — sometimes the agent needs one API doc, sometimes five; sometimes examples, sometimes just types. A fixed workflow template cannot anticipate this.
- Tool-specific. Each AI tool has its own skill/template system. Files work everywhere.
- Decentralized and duplicative. Skills and prompt templates are typically authored by individual teams or companies independently, leading to redundant efforts across the ecosystem. Multiple organizations end up writing overlapping instructions for the same library, each with slightly different interpretations and quality levels. Worse, these community-maintained templates have no built-in mechanism to stay synchronized with the library they describe — when the library ships a new version with API changes, the scattered skills and templates across dozens of teams go stale silently. Co-packaged documentation eliminates this duplication by placing the authoritative instructions at the source: the library author maintains one set of docs, and every consumer receives it automatically on install.
- Probabilistic retrieval is not engineering. Documentation retrieval should not be left to probability. When an agent fails to trigger a skill, or uses it incorrectly, the resulting error compounds: the agent generates incorrect code, the developer corrects it, the agent misinterprets the correction, and the cycle spirals. Co-packaged documentation eliminates the dice roll entirely - the docs are there deterministically, alongside the code they describe. Engineering demands certainty of access, not optimistic hope that correct retrieval will happen.
- Computationally expensive. Running an AI agent is not free. Every failed generation, every correction cycle, every re-prompt consumes tokens — and tokens cost money. Skills that sometimes fire and sometimes don’t create a tax on every interaction: the agent burns compute attempting to determine relevance, potentially retrieves the wrong skill, generates incorrect output, and then must be corrected. In the case of correct skill identification, the probabilistic nature of AI models does not guarantee successful retrieval of relevant documentation creating unnecessary overhead costs. Co-packaged documentation, read directly from the filesystem, is the cheapest possible retrieval mechanism. Why make software engineering more expensive than it needs to be by introducing probabilistic middleware between the agent and the knowledge it requires?
- Distribution and logistics repeat history. Who maintains the skills? Who distributes them? Who ensures they stay current across versions? This is the DefinitelyTyped story replaying in real time. The TypeScript community learned — painfully, over years — that community-maintained type definitions maintained separately from the library source inevitably drift, decay, and fragment. Skills and prompt templates face the identical fate: scattered across repositories, maintained by volunteers with varying commitment, silently stale when the library ships a breaking change. The ecosystem already lived through this with
@types/packages. Co-packaged documentation is the lesson learned: ship it at the source, version it with the code, and eliminate the logistical nightmare of distributed, decoupled maintenance.
6.3 Against Sub-Agents
Spawning a sub-agent to research library documentation adds isolation (the sub-agent can’t see the code being written), inconsistency (parallel sub-agents produce uncoordinated results), and overhead that exceeds the cost of direct file reads.
6.4 The Principle
Don’t add infrastructure when the agent’s existing capabilities — reading files and following references — already solve the problem.
The simplest delivery mechanism is also the most robust: files shipped in the package, with a few lines of rules pointing the agent to them.
7. The Economics of Co-Packaged Documentation
7.1 Cost to Library Authors
The incremental cost of co-packaging documentation is low:
- If documentation already exists (which it does for any established library), the work involves building a transformation pipeline that converts existing docs into AI-friendly structured markdown and integrating it into the build process. This is a one-time investment — typically days to weeks of engineering effort, depending on how structured and consistent the existing documentation is. Libraries with well-organized docs (e.g., generated from JSDoc/Javadoc/Sphinx with consistent templates) will find this straightforward. Libraries with documentation scattered across READMEs, wikis, blog posts, and inline comments will face a harder path, potentially requiring documentation restructuring before the AI transformation pipeline can be effective.
- Ongoing maintenance is extremely low if not zero — if the pipeline generates from the existing documentation source of truth. The AI docs update automatically when the human docs update.
- Package size increase is modest. Markdown is lightweight. The entire Design Factory
.ai/folder — covering 56 components with APIs, examples, guidelines, and demos — compresses to a fraction of the size of a typicalnode_modulestree. For most libraries, the documentation would add less than the size of a single source map file.
For libraries that lack structured documentation entirely, adopting this standard may serve as a catalyst for improving documentation overall — a secondary benefit that accrues to human consumers as well.
7.2 Value to Consumers
The value is disproportionately large:
- Correct code on first generation. The agent uses the right APIs, the right patterns, the right tokens — immediately.
- Version-matched documentation. The docs are always for the exact version installed. No more “this example is for v3 but you have v4” failures.
- Offline capability. Works in air-gapped environments, CI pipelines, corporate networks with restricted internet access.
- Zero configuration. No MCP servers to set up, no API keys to configure, no documentation URLs to bookmark.
7.3 Value to the Ecosystem
At ecosystem scale, co-packaged documentation creates a virtuous cycle:
- Libraries that ship AI-optimized docs produce better AI-generated code.
- Developers prefer libraries that work well with AI assistants.
- Library maintainers have a competitive incentive to ship AI-optimized docs.
- The overall quality of AI-assisted development rises.
This is a classic network effect. The more libraries that adopt the standard, the more reliable AI-assisted development becomes, the more developers demand it from every library.
8. Proposed Standard: Co-packaged Documentation (CoDoc)
We propose a minimal, ecosystem-agnostic specification for co-packaging documentation with library code.
The key words “MUST”, “MUST NOT”, “SHOULD”, “SHOULD NOT”, and “MAY” in this section are to be interpreted as described in RFC 2119, as clarified by RFC 8174, when, and only when, they appear in all capitals, as shown here.
8.1 Directory Convention
Libraries SHOULD include a docs/ directory at the root of the published package containing structured markdown documentation. The choice of a plain, visible directory over a hidden one (such as the .ai/ directory used by the Design Factory reference implementation, Section 5) is deliberate: the documentation is intended for both human and machine consumers and should be plainly visible in the installed package.
package-root/
├── docs/
│ ├── README.md # Entry point: library overview (see Section 8.2)
│ ├── index.md # What the library provides - describes the docs folder structure
│ ├── getting-started.md # Quick start guide
│ ├── lib-dir/ # Library specific directory
│ │ └── ...
│ └── lib-dir-two/
│ └── ... # Structured documentation files
├── src/ # (or lib/, dist/, etc.)
└── package.json # (or setup.py, pom.xml, Cargo.toml, etc.)
8.2 README File
Every docs/ directory MUST contain a README.md file that serves as the entry point.
This file SHOULD cover the library’s purpose (what it does), scope (what it does and does not cover), intended audience, the package version it describes, and the structure of the docs/ directory. This should allow an AI agent to get an overview of the library and reassure itself that it is indeed in the correct directory for the task it is trying to complete.
8.3 Index File
Every docs/ directory MUST contain an index.md file that serves as the catalogue of what the docs/ directory provides.
This file SHOULD:
- List relevant library’s directory files and folders (structure of
docs/directory) - Provide brief descriptions for each
- Link to detailed documentation files using relative paths
The index file enables the agent to discover what’s available and selectively read only what’s relevant.
8.4 Documentation Files
Documentation files SHOULD be plain markdown (UTF-8 encoded) and SHOULD follow these guidelines:
- One concern per file. Separate API reference, usage examples, and guidelines into distinct files. This enables selective reading within context window constraints.
- Machine-parseable structure. Use consistent heading levels, code blocks with language tags, and markdown tables for structured data.
- Working code examples. Include complete, copy-pasteable code snippets — not fragments that require surrounding context to compile.
- Version-awareness. If behavior differs across versions, document the current version only. The docs ship with the version they describe.
8.4.1 Forward-Compatibility Provision
Markdown is the recommended baseline format because current-generation LLMs process it natively and it requires no parsing infrastructure. However, agent architectures are evolving rapidly — future agents may prefer embeddings, structured schemas (JSON-LD, OpenAPI fragments), or indexed databases for documentation retrieval.
To accommodate this evolution, the docs/ directory MAY include an optional manifest.json file alongside the markdown content:
{
"codoc_version": "1.0",
"formats": ["markdown"],
"entry_point": "README.md",
"package_name": "my-library",
"package_version": "2.3.1"
}
This manifest serves as a machine-readable metadata layer that future tooling can extend (e.g., adding "formats": ["markdown", "embeddings"] when an embeddings file is included). The key design constraint is additive evolution: new formats and metadata fields can be added without breaking agents that only understand markdown. The markdown files remain the universal baseline; structured formats are optional enhancements. To avoid drift from the package’s own metadata, package_name and package_version SHOULD be generated automatically at build or packaging time rather than maintained by hand (see Section 8.7).
8.5 Discovery Mechanism
The standard defines a two-tier discovery mechanism, ordered by preference:
-
Auto-discovery (primary). AI tools SHOULD automatically detect
docs/directories in installed dependencies. The presence ofdocs/README.mdin a package root is the canonical signal. This requires no action from the consuming developer and no dependency on external conventions. -
Instruction file integration (supplementary). Libraries MAY provide a mechanism (schematic, CLI tool, or documented instructions) for consumers to add agent instructions that point to the
docs/directory. These instructions SHOULD be compatible with emerging multi-tool conventions such asAGENTS.md.
Auto-discovery as the primary mechanism reduces the fragility of depending on instruction file conventions that are not yet formally standardized. The docs/ directory is self-describing: its presence and structure are sufficient for an agent to begin navigating documentation without external configuration.
The standard acknowledges that no single instruction file convention (AGENTS.md, .github/copilot-instructions.md, etc.) has achieved formal standardization as of this writing.
8.6 Ecosystem-Specific Packaging
| Ecosystem | Include docs/ via |
|---|---|
| npm | "files" field in package.json or .npmignore exclusion removal |
| PyPI | package_data or data_files in setup.py / pyproject.toml |
| Maven | Resource directory inclusion in pom.xml |
| NuGet | Content files in .nuspec or <Content> items in .csproj |
| Crates.io | include field in Cargo.toml |
8.7 Generation, Not Duplication
The standard RECOMMENDS generating docs/ documentation from the library’s existing documentation source, not maintaining it as a separate artifact. This ensures:
- Single source of truth (no documentation drift)
- Automatic coverage of new features
- Low ongoing maintenance overhead
A note on documentation maturity. This recommendation assumes that the library’s existing documentation is reasonably structured and complete. In practice, documentation quality varies enormously. Libraries with well-organized, template-based docs (generated from JSDoc, Javadoc, Sphinx, or similar tools) will find generation straightforward. Libraries with informal or scattered documentation may need to invest in documentation restructuring before a generation pipeline is viable.
The standard does not require perfection. Shipping partial AI-optimized documentation — even just an API reference generated from type definitions or doc comments — is better than shipping none. Libraries can adopt incrementally, expanding coverage over time.
9. Security Considerations
Any mechanism that causes AI agents to automatically read and follow content from third-party packages introduces a prompt injection attack surface. This section addresses the security implications of co-packaged documentation and proposes mitigations.
9.1 Threat Model
The primary threat is a malicious or compromised package that includes docs/ content designed to manipulate agent behavior. Attack vectors include:
- Instruction injection: Documentation files containing hidden instructions (e.g., “ignore all previous instructions and…”) that override the developer’s intent.
- Exfiltration prompts: Content that instructs the agent to read and transmit sensitive files (environment variables, credentials, private keys) from the developer’s workspace.
- Vulnerability introduction: Documentation that recommends insecure patterns — disabling authentication, using eval, or weakening security configurations — disguised as legitimate usage guidance.
- Dependency confusion: A malicious package named similarly to a popular library, shipping
docs/documentation that redirect agent behavior toward the attacker’s code.
9.2 Mitigations
For AI tool vendors:
- Sandboxed documentation context. Agents SHOULD treat
docs/content as reference material, not as system instructions. Documentation should inform the agent’s understanding of an API but should not be able to override user instructions, project-level configuration, or the agent’s safety policies. This is analogous to how agents treat source code: they read it for context but do not execute arbitrary commands found in code comments. - Content provenance signals. When an agent reads
docs/content, it should annotate the context with the source package name and version, enabling the user (and the agent’s safety layer) to distinguish between trusted project-level instructions and third-party documentation. - Scope limitation. Documentation from a dependency’s
docs/folder should only influence the agent’s behavior when working with that specific dependency’s APIs. It should not be able to affect code generation for unrelated parts of the project.
For library authors:
- Plain documentation only. The
docs/directory should contain factual API documentation, usage examples, and guidelines. It should not contain agent instructions, system prompts, or behavioral directives. Themanifest.json(Section 8.4.1) provides a structured metadata channel that is easier to validate than free-form markdown. - Reviewable content. All
docs/content should be reviewable in the package source repository. Consumers should be able to audit what documentation a package ships, just as they can audit source code.
For package registries:
- Content scanning. Registries can scan
docs/directories for known prompt injection patterns (instruction overrides, exfiltration attempts) as part of their existing malware detection pipelines. - Signing and provenance. Package signing (npm provenance, Sigstore for PyPI) provides a chain of trust from the library author to the installed content.
9.3 Risk Assessment
The prompt injection risk for co-packaged documentation is real but bounded. It is comparable in nature — though not in severity — to the existing risk of malicious code in dependencies (supply chain attacks). Developers already accept the risk of running third-party code; co-packaged documentation adds a surface for influencing AI-generated code, which is a lower-severity vector than arbitrary code execution.
The mitigations above reduce the risk to an acceptable level when combined with existing supply chain security practices (lockfiles, dependency auditing, package provenance). The standard acknowledges this risk explicitly and recommends that AI tool vendors treat third-party docs/ content with appropriate skepticism — as context, not as commands.
10. Addressing Objections
10.1 “This bloats package sizes.”
Markdown is extremely compact. A comprehensive documentation set for a medium-complexity library (50–100 public APIs) typically compresses to 100–500 KB. For context:
- The average
node_modulesfolder is 200–500 MB. - A single source map file is often 1–5 MB.
- A single high-resolution image in a README is 100–500 KB.
The documentation is a rounding error in package size while providing outsized value.
10.2 “Documentation goes stale.”
Only if maintained separately. The standard explicitly recommends generating AI-friendly docs from the same source that produces human-facing documentation. When the library updates and the human docs update, the AI-optimized docs update in the same build pipeline. Staleness is a process problem, not an architectural one.
10.3 “My library already has good docs on our website.”
Website-hosted documentation is optimized for human consumption: rich formatting, interactive examples, search widgets, navigation sidebars. AI agents have difficulties with this or cannot use any of this. They need plain-text files on the local filesystem. The generation pipeline transforms your existing good documentation into a format the agent can actually access.
10.4 “AI models already know about popular libraries from training data.”
Training data has a cutoff date. Every library release after that date is invisible to the model. Even for well-known libraries, the agent may confuse APIs across versions, hallucinate deprecated methods, or miss new features. Co-packaged docs provide ground truth for the exact version installed.
10.5 “Can’t the AI just read the source code?”
Source code tells the agent what exists but not how to use it correctly. It doesn’t convey design intent, usage guidelines, do/don’t rules, accessibility requirements, or idiomatic patterns. Documentation is the bridge between “what the code does” and “how to use it well.”
10.6 “Won’t agents just get good at browsing the web?”
They likely will. Agent web-browsing capabilities are improving, and this paper assumes they will continue to do so.
But even with perfect browsing, co-packaged documentation wins on properties that browsing can never provide:
- Version match. A browsing agent finds the documentation the library’s website currently hosts, this is typically the latest release. Co-packaged documentation describes the exact version installed in the project, which in practice is often not the latest.
- Offline and air-gapped use. Secure enterprise networks, regulated environments have no reliable internet access. Local files work everywhere the installed package works.
- Latency. Every web fetch is a network round trip involving DNS, TLS, page rendering, and HTML parsing. A local file read is effectively instant and agents perform hundreds of reads per session.
- Determinism. Websites change between requests: content is reorganized, redesigned, A/B tested. The same query can return different content on different days. A co-packaged file is byte-identical for every agent, every session, every time.
- Corporate network restrictions. Documentation portals behind authentication, SSO, or VPNs are unreachable to a browsing agent. Local files have no such gate.
- Anti-bot controls. Even publicly available documentation may sit behind CAPTCHAs, JavaScript challenges, bot detection, or aggressive rate limits. These controls are reasonable defenses for a website but can block or degrade an agent’s access. A local file does not need to prove that its reader is human.
- Security. Every page a browsing agent retrieves is untrusted content the agent must ingest and obey-like text; a hijacked site, an injected instruction, or a lookalike domain can steer code generation. Co-packaged documentation arrives through the same vetted, integrity-checked channel as the code itself, so the trust boundary does not widen. Neither source is risk-free (Section 9 addresses prompt injection in co-packaged docs) but browsing multiplies the attack surface with the entire open web.
Better browsing raises the floor for agents but it does not close the version-match, availability, latency, or determinism gaps, because those are properties of where the documentation lives, not of how capable the reader is. Section 10.3 addresses today’s format mismatch; this objection fails even on tomorrow’s capabilities.
10.7 “Why not point the agent at the git repository?”
Reading the repository or raw files served from a Git host is useful, but the repository is a weaker source of truth than the installed package:
- Repositories drift from releases. The default branch reflects unreleased work in progress. An agent reading it may learn APIs that have not shipped yet, or that shipped with different signatures.
- Tags are not guarantees. Tags can move, be renamed, or be deleted, and they need not correspond exactly to what the registry distributed. What
npm install(orpip install, or the Maven equivalent) placed in the project’s dependencies is the only version that matters, and it is a registry artifact, not a git checkout. - Published registry artifacts are immutable and version-locked. npm, PyPI, and Maven Central do not allow a published version to be overwritten, barring exceptional removals. The documentation inside
example-lib@2.3.1today is the documentation insideexample-lib@2.3.1forever, and it is guaranteed to describe the exact code it accompanies. - Access may require authentication. Private repositories, SSO-gated Git hosts, and rate-limited APIs introduce friction and failure modes. The installed package is already local; permissions were resolved at install time.
The registry artifact is the contract between the library author and the consumer. The repository is the workshop. Agents should read the contract.
10.8 “This doesn’t scale to deep dependency graphs.”
A typical Node.js project has hundreds to thousands of transitive dependencies. If every one ships docs/ documentation, does the agent drown in documentation?
This is a legitimate concern, and the answer is scoped relevance, not exhaustive ingestion. The agent should not read docs/ documentation for all 800 dependencies upfront. Instead:
- On-demand reading. The agent consults a package’s
docs/documentation only when it encounters that package’s API in the code being written or modified. Most transitive dependencies are never directly referenced by application code. - Direct dependencies first. Agent tooling should prioritize
docs/documentation from direct dependencies (listed inpackage.json,requirements.txt, orpom.xml) over transitive ones. - Index-level scanning. The
README.mdentry point is lightweight enough for the agent to scan across multiple packages to determine relevance before reading detailed docs.
In practice, a developer typically interacts directly with 5–20 libraries in a given coding session. The agent needs docs for those libraries, not for the entire dependency tree. This is the same scoping that developers apply naturally — you don’t read the docs for every transitive dependency; you read the docs for the libraries you’re calling.
10.9 “The generation pipeline is too complex for most libraries.”
The cost of building a documentation generation pipeline is real (see Section 7.1). However, the standard is designed for incremental adoption:
- Tier 1 (minimal effort): Ship your existing README and API reference as
docs/README.md. This requires no pipeline — just copying files into the package. - Tier 2 (moderate effort): Generate structured markdown from existing doc comments (JSDoc, Javadoc, docstrings) using widely available tools. This is a one-time build step.
- Tier 3 (full investment): Build a transformation pipeline from your documentation source (Sphinx, Docusaurus, custom CMS) to structured
docs/output. This is the aspirational target but not the entry bar.
The ecosystem can support adoption at all tiers. Even Tier 1 — a well-written README.md file covering the library’s purpose, scope, and primary APIs (Section 8.2) — provides meaningful value over no documentation at all.
11. A Path Forward
11.1 Governance and Standardization
This paper proposes a new industry standard. For the standard to achieve the ecosystem-wide adoption it aspires to, it needs a governance path. We propose the following trajectory:
- Community RFC phase (current). This paper serves as the initial request for comments. We invite feedback, critiques, and counter-proposals from library authors, AI tool vendors, and the developer community.
- Working group formation. Interested parties form a cross-ecosystem working group to refine the specification, address edge cases, produce a formal specification document, and develop the public CoDocBench evaluation suite (Section 2.3) for validating the standard’s effectiveness. Natural homes for this working group include the OpenJS Foundation (for npm), the Python Packaging Authority (for PyPI), or a cross-ecosystem body.
- Tool vendor alignment. As the specification stabilizes, AI tool vendors implement auto-discovery of
docs/directories, reducing the dependency on instruction file conventions and providing the agent-side infrastructure for the standard. - Registry integration. Package registries adopt metadata signals for co-packaged documentation, providing visibility and incentives for adoption.
The absence of a governance body is a known weakness at this stage. We address it directly rather than assuming adoption will occur organically.
11.2 For Library Authors
- Start with what you have. If you have existing documentation (and you likely do), build a transformation pipeline that converts it to AI-friendly markdown.
- Add a
docs/directory to your package. Include aREADME.mdentry point and structured documentation files. - Integrate into your build. Make documentation generation a build step, not a manual process. When you release a new version, the AI-optimized docs update automatically.
- Provide agent instructions. Offer a template or tool for consumers to add
docs/references to their AI tool configuration.
11.3 For Package Registries
Package registries (npm, PyPI, Maven Central) can accelerate adoption by:
- Recognizing the
docs/convention. Display a badge or indicator when a package includes AI documentation. - Including documentation quality in package rankings. Just as registries surface type definitions and test coverage, they could surface AI-optimized documentation completeness.
- Providing guidelines and tooling. Publish documentation transformation tools that library authors can adopt.
11.4 For AI Tool Vendors
AI coding assistants can support the standard by:
- Auto-discovering
docs/directories in installed dependencies when no explicit instructions are configured. This is the single most impactful action tool vendors can take to reduce adoption friction. - Sandboxing third-party documentation context. Treat
docs/content as reference material, not as system instructions, to mitigate prompt injection risks (see Section 9). - Indexing co-packaged documentation for faster retrieval during code generation.
- Preferring co-packaged docs over training data when both are available, since the packaged docs are version-matched and authoritative.
11.5 For the Developer Community
Developers can drive adoption by:
- Requesting AI-optimized documentation from library maintainers, just as the community once requested TypeScript type definitions.
- Contributing documentation pipelines to open-source libraries.
- Sharing transformation tooling across ecosystems.
12. Historical Precedent: The TypeScript Analogy
The co-packaged documentation proposal follows a pattern the ecosystem has seen before: the DefinitelyTyped trajectory.
In TypeScript’s early years, most npm packages shipped without type definitions. The community created DefinitelyTyped — a separate repository of type definitions maintained by volunteers. Developers installed types separately (npm install --save-dev @types/lodash). This worked, but:
- Types drifted from the actual library code.
- Maintenance was a community burden.
- Coverage was incomplete.
Over time, library authors began shipping types directly in their packages ("types" field in package.json). Today, first-party types are the expectation, and DefinitelyTyped is a fallback for legacy packages.
AI documentation is on the same trajectory. Today, AI-optimized docs are rare and, where they exist, are maintained separately. Tomorrow, they will be expected to ship with the package. The Think Inside the Box Principle simply names this inevitable evolution and proposes a standard to accelerate it.
13. Conclusion
The separation of code and documentation made sense when the documentation consumer was a human with a web browser. Keeping this as the de facto standard no longer makes sense when the main consumer is an AI agent with a file reader.
The Think Inside the Box Principle — everything in one box — is a proven model for product delivery. Applied to software libraries, it means shipping structured, AI-optimized documentation alongside the code in the same package artifact. The benefits are significant (more accurate AI-generated code, version-matched docs, offline capability) and the costs are negligible (markdown is small, generation pipelines are automatable).
This paper claims that the direction towards shipping code and documentation in one box is the right path forward, that the timing is urgent, and the entry bar is low enough for immediate adoption. Design Factory demonstrates that the approach works in their preliminary testing today and are ready to continue shipping co-packaged AI-optimized documentation in the near future. The patterns it established — a directory with an entry point, structured markdown files, generated-not-maintained documentation, and layered discovery mechanisms (standard agent instructions) — generalize cleanly across ecosystems.
We are at an inflection point. The ecosystem moved from “types are someone else’s problem” to “types ship with the package.” The same shift is beginning for AI-optimized documentation. The question is not whether libraries will ship AI-optimized docs, but how quickly the ecosystem converges on a standard for doing so.
This paper proposes that standard. We invite library authors, package registry maintainers, AI tool vendors, and the developer community to adopt, refine, and propagate it.
Ship the Docs in the Box and remember to Think Inside the Box
References
- Design Factory AI Documentation Architecture — Internal technical documentation, 2026. To be made publicly available as part of Design Factory’s open-sourcing [TODO: public repository link when available].
- Model Context Protocol (MCP) Specification — Anthropic, 2025.
AGENTS.mdConvention — Emerging multi-tool convention for AI coding agent instructions.- DefinitelyTyped — Community-maintained TypeScript type definitions
- OWASP LLM Top 10 — Prompt Injection risks in LLM applications
- npm Provenance — Supply chain security for npm packages
- Sigstore — Software signing and transparency for open source
- OpenJS Foundation
- Python Packaging Authority (PyPA)
- S. Bradner — RFC 2119: Key words for use in RFCs to Indicate Requirement Levels
- B. Leiba — RFC 8174: Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words
llms.txt— A proposal to standardize on using an /llms.txt file to provide LLM-friendly website content (Jeremy Howard, 2024)- Context7 — Up-to-date Code Docs For Any Prompt
- DeepWiki — Cognition AI: AI-generated documentation wikis for public code repositories
- GitHub Copilot — Adding repository custom instructions for GitHub Copilot
- Cursor — Cursor Rules: project-level instructions for the Cursor AI code editor
This white paper is released for public discussion. We welcome contributions, critiques, and adoption reports from the software engineering community.