Think Inside the Box for Software Libraries

A White Paper on Co-Packaging Code and AI-Optimized Documentation as an Industry Standard

Authors: AI Documentation Working Group - Andres IGEA OVIEDO (andres.igeaoviedo@amadeus.com), Andrea SCORTI (andrea.scorti@amadeus.com)
Date: 29th of September 2026
Version: 1.1


Abstract

The rapid adoption of AI coding assistants — GitHub Copilot, Claude, Codex, and others — has fundamentally changed how developers consume software libraries. Yet the ecosystem’s documentation infrastructure remains anchored in a pre-AI paradigm: code ships in a package, documentation lives elsewhere. This disconnect forces AI agents to operate without the contextual knowledge they need, producing incorrect API usage, outdated patterns, and hallucinated interfaces.

This paper argues for a new industry standard: co-packaged documentation (CoDoc) — the practice of shipping structured, AI-optimized documentation alongside library code within the same distributable artifact. We call this the Think Inside the Box Principle: everything the consumer needs arrives in one box.

We present a reference implementation from the Design Factory design system, demonstrate that the approach works across package ecosystems (npm, pip, Maven), and propose a minimal specification that library authors can adopt today.

We also address critical concerns around security (prompt injection via co-packaged content), forward-compatibility with evolving agent architectures, scalability across deep dependency graphs, and the practical challenges of documentation generation pipelines.


1. Introduction: The AI-Driven Development Shift

Software development is undergoing its most significant workflow transformation since the introduction of the integrated development environment. AI coding assistants are no longer novelties — they are daily tools for millions of developers. GitHub reports that Copilot generates a significant share of new code in enabled repositories. Anthropic’s Claude, OpenAI’s Codex, and a growing ecosystem of AI agents are being embedded into every stage of the development lifecycle.

These agents share a common trait: they are only as effective as the context they can access. An AI assistant writing React code with full access to the React documentation produces dramatically better output than one relying solely on its training data. Training data goes stale; documentation is versioned and current.

Yet today, when a developer installs a library — whether via npm install, pip install, or a Maven dependency — the AI assistant working in that project receives the library’s code but almost never its documentation. The documentation exists, but it lives on a website, behind a search engine, in a format optimized for human browsing rather than machine consumption.

This is the gap we propose to close.


2. The Current State: A Separated World

2.1 How Libraries Ship Today

Across every major package ecosystem, the convention is the same: the package contains code; the documentation is hosted elsewhere.

Ecosystem Package Contents Documentation Location
npm JavaScript/TypeScript source, package.json, README Dedicated docs site, GitHub wiki, or MDN
PyPI Python source, setup.py/pyproject.toml, README Read the Docs, Sphinx-hosted site, or GitHub pages
Maven Compiled JARs, POM metadata Javadoc on Maven Central, project wiki, or vendor site
NuGet .NET assemblies, XML doc comments Microsoft Learn, GitHub, or custom docs portals
Crates.io Rust source, Cargo.toml docs.rs (auto-generated)

The README file — the one piece of documentation that does travel with the package — is typically a brief overview: a logo, a one-liner description, installation instructions, and a link to “full documentation” hosted online.

This separation made sense in a human-centric world. Developers browse documentation in a web browser with search, navigation, and rich formatting. Shipping a full documentation site inside every node_modules dependency would have been wasteful.

But the consumer has changed. The primary documentation consumer is increasingly an AI agent running locally in the developer’s IDE, and the agent’s capabilities in browsing websites is not as robust as the capabilities it has for searching, navigating and reading files in a local workspace.

2.2 What AI Agents Can and Cannot Do

Modern AI coding assistants operate within a file-based context window. They can:

They typically have difficulty with, cannot or should not need to:

When an AI agent encounters an unfamiliar library in a project’s dependencies, it has three options:

  1. Rely on training data — which may be months or years out of date, missing recent API changes, deprecations, or new features.
  2. Search the web — which adds latency, requires network access (often unavailable in corporate environments), and returns HTML pages optimized for human reading.
  3. Read local files — fast, always available, version-matched to the installed library, and optimized for machine consumption.

Option 3 is clearly superior. But today, there is almost nothing to read.

2.3 The Consequences of the Gap

The documentation gap produces observable failures in AI-assisted development:

These are not edge cases. They are the daily experience of developers using AI tools with any library that has evolved since the agent’s training cutoff. Preliminary observations from Design Factory’s internal adoption suggest that co-packaged documentation significantly reduces these failures, though rigorous benchmarking across diverse libraries remains an open research direction.

To close this evidence gap, the authors intend to develop CoDocBench, a public evaluation suite measuring AI-agent performance on library-usage tasks (for example, API correctness and deprecated-pattern rates) with and without co-packaged documentation, across multiple libraries and package ecosystems. CoDocBench is intended to provide measurable success criteria for the standard proposed in this paper and to invite external validation of its claims.


CoDoc does not exist in a vacuum. Several existing conventions already address fragments of the problem this paper targets. This section surveys the closest adjacent work. Analysing this work reveals an emerging pattern: existing efforts focus on the format problem (how to write documentation for AI consumption) or on providing hosted retrieval (serving documentation from a registry or third party), but none solve the delivery and logistics problem: getting version-matched documentation onto the agent’s filesystem, inside the package artifact.

3.1 llms.txt: The Right Format in the Wrong Location

The closest existing convention is llms.txt, a proposal that websites publish a markdown file /llms.txt at the site’s root or at any path within it, with an optional /llms-full.txt containing the entire collection of the site’s core content so that AI agents can consume its content without parsing HTML. It has seen adoption across some documentation sites and developer tools.

llms.txt gets the format right: curated markdown, an index enabling selective reading, prose that can be mixed with code examples. These are some of the same properties Section 8.4 specifies for CoDoc documentation files, and the convergence is encouraging. The ecosystem is independently arriving at markdown as the canonical AI-consumable format.

Where llms.txt differs is location, and location is precisely the problem. An llms.txt file lives on a website, so the versioning, offline, and determinism arguments of Section 10.6 apply to it directly:

The two conventions are complementary: llms.txt answers “how should a website serve AI agents?”, while CoDoc answers “how should a package serve AI agents?”.

3.2 Hosted Documentation Intermediaries

A growing category of services such as Context7, DeepWiki, and similar documentation-serving layers index library documentation or source repositories and serve the result to AI agents, typically through MCP tools or web interfaces. These services demonstrate real demand for agent-consumable documentation and are genuinely useful stopgaps.

In the framing of this paper, however, they are probabilistic middleware between the agent and the knowledge it requires (Section 6.2):

These services are a valuable bridge for the long tail of libraries that ship no agent-consumable documentation at all. But a bridge is not a foundation. The authoritative, version-matched source should ship from the library author, inside the package.

3.3 Agent Instruction Conventions: Complementary Infrastructure

A parallel set of conventions has emerged for instructing agents how to behave inside a project: AGENTS.md, .github/copilot-instructions.md, Cursor rules, and their tool-specific variants. CoDoc does not compete with these conventions; it gives them something stable to point at.

Instruction files are pointers; co-packaged documentation is the pointed-at content. Without CoDoc, an instruction file that wants to ground the agent in a library’s API can only point at a URL — inheriting the version-mismatch, availability, and determinism problems of Section 10.6. With CoDoc, the same instruction file can point out to the agent that all software libraries have co-packaged documentation alongside the library code within the same distributable artifact. CoDoc scales elegantly and avoids the hardcoding of URLs anywhere in a software engineering project. This is exactly the two-tier discovery model defined in Section 8.5: auto-discovery of docs/ directories as the primary mechanism, instruction-file integration as the supplementary channel.

3.4 Summary

Convention AI-optimized format Version-matched Offline-capable Maintained at the source
llms.txt Yes Not guaranteed No Yes
Hosted intermediaries (Context7, DeepWiki) Yes Sometimes (indexed version) No No (third party)
Instruction files (AGENTS.md, etc.) N/A — instructions, not documentation N/A Yes N/A (consumer-authored)
CoDoc Yes Yes Yes Yes

No existing convention delivers version-matched, AI-optimized documentation inside the package artifact. This specific combination is CoDoc’s contribution.


4. The Think Inside the Box Principle

Some physical goods companies’ success is built on a deceptively simple idea: everything you need comes in one box. For instance, when you purchase a bookcase, the box contains the panels, the shelves, the screws, the dowels, the cam locks, and — critically — the assembly instructions. You do not need to visit a website to look up how to assemble it. You do not need to search YouTube for a tutorial. The instructions are right there, designed to be consumed alongside the product, version-matched and complete.

This principle has a direct analog in software:

Think Inside the Box for Software Libraries: A package should contain everything an agent needs to correctly use the library — the code and its documentation — in the same distributable artifact.

The analogy is precise:

Physical goods company Software Library
Product (panels, shelves) Source code, compiled artifacts
Assembly and usage instructions API docs, usage guides, examples
The box The package (npm tarball, wheel, JAR)
The customer The AI coding agent
The store The package registry (npm, PyPI, Maven Central)

Just as instructions from newly purchased furniture are designed for the person assembling it and using it — not the designer who created it — co-packaged documentation should be designed for the agent consuming the library, not the author who wrote it.

The key insight is that the consumer of library documentation is changing. Historically, documentation was written for human developers who read it in a browser. Increasingly, documentation must also serve AI agents that read it as local files. The Think Inside the Box Principle asks library authors to serve this new consumer by including documentation in the delivery, not just linking to it.


5. Design Factory: A Reference Implementation

The Design Factory design system provides a concrete implementation of co-packaged documentation. Analyzing its architecture reveals patterns that generalize across ecosystems.

5.1 What Ships in the Package

Design Factory publishes the following structure inside its npm package:

@design-factory/design-factory/
├── .ai/
│   ├── index.md                          # Component catalog (56 components)
│   ├── design-tokens.md                  # 4-tier token reference
│   ├── foundations.md                    # Typography, spacing, grid
│   └── docs/
│       ├── components/{slug}/
│       │   ├── api.md                    # Selectors, inputs, outputs
│       │   ├── overview.md               # Component overview
│       │   ├── examples.md               # Usage examples with do/don't
│       │   ├── guidelines.md             # Usage guidelines
│       │   └── accessibility.md          # Accessibility guidance
│       └── demos/{slug}/{variant}/
│           ├── component-variant.ts      # Working demo source
│           └── component-variant.html    # Demo template

This is the documentation, versioned and shipped with the code. When a developer runs npm install @design-factory/design-factory, they receive not just the Angular components but a complete, structured knowledge base that any AI agent can navigate.

A note on the .ai/ directory name. Design Factory’s first implementation predates the current CoDoc standard proposed in Section 8, which specifies a plain docs/ directory instead of a hidden .ai/ directory. Early feedback on the .ai/ convention recommended a location that is visible and accessible to both humans and AI agents: hidden dot-directories are easy to overlook in file explorers, and some tools exclude them from search and indexing by default. Design Factory will update its package structure to conform to the standard proposed in this paper, and is in the process of becoming fully open-source, which will make its documentation generation pipeline available as a public reference implementation.

5.2 How the Agent Uses It

The integration works through two complementary discovery mechanisms:

Auto-discovery (preferred). AI tool vendors are encouraged to automatically detect .ai/ directories in installed dependencies (see Section 11.4). When an agent encounters an unfamiliar API, it scans for a .ai/index.md in the relevant package and navigates from there. This requires zero configuration from the consuming developer.

Explicit instructions (supplementary). For tools that do not yet support auto-discovery, a standard instructions file (AGENTS.md) in the project root can direct the agent to .ai/ folders in specific dependencies:

  1. The AI agent reads AGENTS.md in the project root.
  2. The rules direct it to the .ai/ folder in the installed package.
  3. The agent reads index.md to understand what components exist.
  4. For a specific task (e.g., building a form with a datepicker), the agent selectively reads only the relevant docs — API, examples, guidelines.
  5. The agent produces code using the correct components, tokens, and patterns.

No web browsing. No external tools. No MCP servers. No infrastructure. Just files.

5.3 Key Design Decisions That Generalize

Several decisions in the Design Factory implementation reflect principles that apply to any library:

Plain markdown over structured formats. LLMs are trained on vast amounts of markdown. It’s the format they understand best, allows mixing prose with code examples, and is readable by both humans and machines. JSON schemas or YAML configs would require parsing layers with no added benefit.

Dynamic discovery over front-loaded injection. Rather than injecting all documentation into the AI prompt, the system provides an index file. The agent reads the index, identifies what’s relevant, and selectively reads only the docs it needs. This respects context window limits and mirrors how a competent developer navigates documentation.

Single source of truth. The shipped documentation is generated from the same sources that power the human-facing documentation portal. There is no separate “AI version” to maintain. This eliminates documentation drift — when the library updates, the docs update with it.


6. Why Static Files Are the Pragmatic Starting Point

A natural objection is: “Wouldn’t it be better to expose documentation through an MCP server, a dedicated tool, or a documentation API?” We argue that static files should be the foundation, and the reasoning is pragmatic.

6.1 Against MCP Servers for Documentation

The Model Context Protocol (MCP) allows AI agents to call external tools during a conversation. An MCP server for a library could offer tools like get_component_docs(name) or search_api(query). However we have to highlight the following disadvantages:

6.2 Against Skills and Prompt Templates

Skills (pre-written prompt expansions) inject multi-step workflows when triggered. A “use library X” skill could instruct the agent to read docs in a specific order. However:

6.3 Against Sub-Agents

Spawning a sub-agent to research library documentation adds isolation (the sub-agent can’t see the code being written), inconsistency (parallel sub-agents produce uncoordinated results), and overhead that exceeds the cost of direct file reads.

6.4 The Principle

Don’t add infrastructure when the agent’s existing capabilities — reading files and following references — already solve the problem.

The simplest delivery mechanism is also the most robust: files shipped in the package, with a few lines of rules pointing the agent to them.


7. The Economics of Co-Packaged Documentation

7.1 Cost to Library Authors

The incremental cost of co-packaging documentation is low:

For libraries that lack structured documentation entirely, adopting this standard may serve as a catalyst for improving documentation overall — a secondary benefit that accrues to human consumers as well.

7.2 Value to Consumers

The value is disproportionately large:

7.3 Value to the Ecosystem

At ecosystem scale, co-packaged documentation creates a virtuous cycle:

  1. Libraries that ship AI-optimized docs produce better AI-generated code.
  2. Developers prefer libraries that work well with AI assistants.
  3. Library maintainers have a competitive incentive to ship AI-optimized docs.
  4. The overall quality of AI-assisted development rises.

This is a classic network effect. The more libraries that adopt the standard, the more reliable AI-assisted development becomes, the more developers demand it from every library.


8. Proposed Standard: Co-packaged Documentation (CoDoc)

We propose a minimal, ecosystem-agnostic specification for co-packaging documentation with library code.

The key words “MUST”, “MUST NOT”, “SHOULD”, “SHOULD NOT”, and “MAY” in this section are to be interpreted as described in RFC 2119, as clarified by RFC 8174, when, and only when, they appear in all capitals, as shown here.

8.1 Directory Convention

Libraries SHOULD include a docs/ directory at the root of the published package containing structured markdown documentation. The choice of a plain, visible directory over a hidden one (such as the .ai/ directory used by the Design Factory reference implementation, Section 5) is deliberate: the documentation is intended for both human and machine consumers and should be plainly visible in the installed package.

package-root/
├── docs/
│   ├── README.md           # Entry point: library overview (see Section 8.2)
│   ├── index.md            # What the library provides - describes the docs folder structure
│   ├── getting-started.md  # Quick start guide
│   ├── lib-dir/            # Library specific directory
│   │   └── ...
│   └── lib-dir-two/
│       └── ...             # Structured documentation files
├── src/                    # (or lib/, dist/, etc.)
└── package.json            # (or setup.py, pom.xml, Cargo.toml, etc.)

8.2 README File

Every docs/ directory MUST contain a README.md file that serves as the entry point.

This file SHOULD cover the library’s purpose (what it does), scope (what it does and does not cover), intended audience, the package version it describes, and the structure of the docs/ directory. This should allow an AI agent to get an overview of the library and reassure itself that it is indeed in the correct directory for the task it is trying to complete.

8.3 Index File

Every docs/ directory MUST contain an index.md file that serves as the catalogue of what the docs/ directory provides.

This file SHOULD:

The index file enables the agent to discover what’s available and selectively read only what’s relevant.

8.4 Documentation Files

Documentation files SHOULD be plain markdown (UTF-8 encoded) and SHOULD follow these guidelines:

8.4.1 Forward-Compatibility Provision

Markdown is the recommended baseline format because current-generation LLMs process it natively and it requires no parsing infrastructure. However, agent architectures are evolving rapidly — future agents may prefer embeddings, structured schemas (JSON-LD, OpenAPI fragments), or indexed databases for documentation retrieval.

To accommodate this evolution, the docs/ directory MAY include an optional manifest.json file alongside the markdown content:

{
  "codoc_version": "1.0",
  "formats": ["markdown"],
  "entry_point": "README.md",
  "package_name": "my-library",
  "package_version": "2.3.1"
}

This manifest serves as a machine-readable metadata layer that future tooling can extend (e.g., adding "formats": ["markdown", "embeddings"] when an embeddings file is included). The key design constraint is additive evolution: new formats and metadata fields can be added without breaking agents that only understand markdown. The markdown files remain the universal baseline; structured formats are optional enhancements. To avoid drift from the package’s own metadata, package_name and package_version SHOULD be generated automatically at build or packaging time rather than maintained by hand (see Section 8.7).

8.5 Discovery Mechanism

The standard defines a two-tier discovery mechanism, ordered by preference:

  1. Auto-discovery (primary). AI tools SHOULD automatically detect docs/ directories in installed dependencies. The presence of docs/README.md in a package root is the canonical signal. This requires no action from the consuming developer and no dependency on external conventions.

  2. Instruction file integration (supplementary). Libraries MAY provide a mechanism (schematic, CLI tool, or documented instructions) for consumers to add agent instructions that point to the docs/ directory. These instructions SHOULD be compatible with emerging multi-tool conventions such as AGENTS.md.

Auto-discovery as the primary mechanism reduces the fragility of depending on instruction file conventions that are not yet formally standardized. The docs/ directory is self-describing: its presence and structure are sufficient for an agent to begin navigating documentation without external configuration.

The standard acknowledges that no single instruction file convention (AGENTS.md, .github/copilot-instructions.md, etc.) has achieved formal standardization as of this writing.

8.6 Ecosystem-Specific Packaging

Ecosystem Include docs/ via
npm "files" field in package.json or .npmignore exclusion removal
PyPI package_data or data_files in setup.py / pyproject.toml
Maven Resource directory inclusion in pom.xml
NuGet Content files in .nuspec or <Content> items in .csproj
Crates.io include field in Cargo.toml

8.7 Generation, Not Duplication

The standard RECOMMENDS generating docs/ documentation from the library’s existing documentation source, not maintaining it as a separate artifact. This ensures:

A note on documentation maturity. This recommendation assumes that the library’s existing documentation is reasonably structured and complete. In practice, documentation quality varies enormously. Libraries with well-organized, template-based docs (generated from JSDoc, Javadoc, Sphinx, or similar tools) will find generation straightforward. Libraries with informal or scattered documentation may need to invest in documentation restructuring before a generation pipeline is viable.

The standard does not require perfection. Shipping partial AI-optimized documentation — even just an API reference generated from type definitions or doc comments — is better than shipping none. Libraries can adopt incrementally, expanding coverage over time.


9. Security Considerations

Any mechanism that causes AI agents to automatically read and follow content from third-party packages introduces a prompt injection attack surface. This section addresses the security implications of co-packaged documentation and proposes mitigations.

9.1 Threat Model

The primary threat is a malicious or compromised package that includes docs/ content designed to manipulate agent behavior. Attack vectors include:

9.2 Mitigations

For AI tool vendors:

For library authors:

For package registries:

9.3 Risk Assessment

The prompt injection risk for co-packaged documentation is real but bounded. It is comparable in nature — though not in severity — to the existing risk of malicious code in dependencies (supply chain attacks). Developers already accept the risk of running third-party code; co-packaged documentation adds a surface for influencing AI-generated code, which is a lower-severity vector than arbitrary code execution.

The mitigations above reduce the risk to an acceptable level when combined with existing supply chain security practices (lockfiles, dependency auditing, package provenance). The standard acknowledges this risk explicitly and recommends that AI tool vendors treat third-party docs/ content with appropriate skepticism — as context, not as commands.


10. Addressing Objections

10.1 “This bloats package sizes.”

Markdown is extremely compact. A comprehensive documentation set for a medium-complexity library (50–100 public APIs) typically compresses to 100–500 KB. For context:

The documentation is a rounding error in package size while providing outsized value.

10.2 “Documentation goes stale.”

Only if maintained separately. The standard explicitly recommends generating AI-friendly docs from the same source that produces human-facing documentation. When the library updates and the human docs update, the AI-optimized docs update in the same build pipeline. Staleness is a process problem, not an architectural one.

10.3 “My library already has good docs on our website.”

Website-hosted documentation is optimized for human consumption: rich formatting, interactive examples, search widgets, navigation sidebars. AI agents have difficulties with this or cannot use any of this. They need plain-text files on the local filesystem. The generation pipeline transforms your existing good documentation into a format the agent can actually access.

Training data has a cutoff date. Every library release after that date is invisible to the model. Even for well-known libraries, the agent may confuse APIs across versions, hallucinate deprecated methods, or miss new features. Co-packaged docs provide ground truth for the exact version installed.

10.5 “Can’t the AI just read the source code?”

Source code tells the agent what exists but not how to use it correctly. It doesn’t convey design intent, usage guidelines, do/don’t rules, accessibility requirements, or idiomatic patterns. Documentation is the bridge between “what the code does” and “how to use it well.”

10.6 “Won’t agents just get good at browsing the web?”

They likely will. Agent web-browsing capabilities are improving, and this paper assumes they will continue to do so.

But even with perfect browsing, co-packaged documentation wins on properties that browsing can never provide:

Better browsing raises the floor for agents but it does not close the version-match, availability, latency, or determinism gaps, because those are properties of where the documentation lives, not of how capable the reader is. Section 10.3 addresses today’s format mismatch; this objection fails even on tomorrow’s capabilities.

10.7 “Why not point the agent at the git repository?”

Reading the repository or raw files served from a Git host is useful, but the repository is a weaker source of truth than the installed package:

The registry artifact is the contract between the library author and the consumer. The repository is the workshop. Agents should read the contract.

10.8 “This doesn’t scale to deep dependency graphs.”

A typical Node.js project has hundreds to thousands of transitive dependencies. If every one ships docs/ documentation, does the agent drown in documentation?

This is a legitimate concern, and the answer is scoped relevance, not exhaustive ingestion. The agent should not read docs/ documentation for all 800 dependencies upfront. Instead:

In practice, a developer typically interacts directly with 5–20 libraries in a given coding session. The agent needs docs for those libraries, not for the entire dependency tree. This is the same scoping that developers apply naturally — you don’t read the docs for every transitive dependency; you read the docs for the libraries you’re calling.

10.9 “The generation pipeline is too complex for most libraries.”

The cost of building a documentation generation pipeline is real (see Section 7.1). However, the standard is designed for incremental adoption:

The ecosystem can support adoption at all tiers. Even Tier 1 — a well-written README.md file covering the library’s purpose, scope, and primary APIs (Section 8.2) — provides meaningful value over no documentation at all.


11. A Path Forward

11.1 Governance and Standardization

This paper proposes a new industry standard. For the standard to achieve the ecosystem-wide adoption it aspires to, it needs a governance path. We propose the following trajectory:

  1. Community RFC phase (current). This paper serves as the initial request for comments. We invite feedback, critiques, and counter-proposals from library authors, AI tool vendors, and the developer community.
  2. Working group formation. Interested parties form a cross-ecosystem working group to refine the specification, address edge cases, produce a formal specification document, and develop the public CoDocBench evaluation suite (Section 2.3) for validating the standard’s effectiveness. Natural homes for this working group include the OpenJS Foundation (for npm), the Python Packaging Authority (for PyPI), or a cross-ecosystem body.
  3. Tool vendor alignment. As the specification stabilizes, AI tool vendors implement auto-discovery of docs/ directories, reducing the dependency on instruction file conventions and providing the agent-side infrastructure for the standard.
  4. Registry integration. Package registries adopt metadata signals for co-packaged documentation, providing visibility and incentives for adoption.

The absence of a governance body is a known weakness at this stage. We address it directly rather than assuming adoption will occur organically.

11.2 For Library Authors

  1. Start with what you have. If you have existing documentation (and you likely do), build a transformation pipeline that converts it to AI-friendly markdown.
  2. Add a docs/ directory to your package. Include a README.md entry point and structured documentation files.
  3. Integrate into your build. Make documentation generation a build step, not a manual process. When you release a new version, the AI-optimized docs update automatically.
  4. Provide agent instructions. Offer a template or tool for consumers to add docs/ references to their AI tool configuration.

11.3 For Package Registries

Package registries (npm, PyPI, Maven Central) can accelerate adoption by:

11.4 For AI Tool Vendors

AI coding assistants can support the standard by:

11.5 For the Developer Community

Developers can drive adoption by:


12. Historical Precedent: The TypeScript Analogy

The co-packaged documentation proposal follows a pattern the ecosystem has seen before: the DefinitelyTyped trajectory.

In TypeScript’s early years, most npm packages shipped without type definitions. The community created DefinitelyTyped — a separate repository of type definitions maintained by volunteers. Developers installed types separately (npm install --save-dev @types/lodash). This worked, but:

Over time, library authors began shipping types directly in their packages ("types" field in package.json). Today, first-party types are the expectation, and DefinitelyTyped is a fallback for legacy packages.

AI documentation is on the same trajectory. Today, AI-optimized docs are rare and, where they exist, are maintained separately. Tomorrow, they will be expected to ship with the package. The Think Inside the Box Principle simply names this inevitable evolution and proposes a standard to accelerate it.


13. Conclusion

The separation of code and documentation made sense when the documentation consumer was a human with a web browser. Keeping this as the de facto standard no longer makes sense when the main consumer is an AI agent with a file reader.

The Think Inside the Box Principle — everything in one box — is a proven model for product delivery. Applied to software libraries, it means shipping structured, AI-optimized documentation alongside the code in the same package artifact. The benefits are significant (more accurate AI-generated code, version-matched docs, offline capability) and the costs are negligible (markdown is small, generation pipelines are automatable).

This paper claims that the direction towards shipping code and documentation in one box is the right path forward, that the timing is urgent, and the entry bar is low enough for immediate adoption. Design Factory demonstrates that the approach works in their preliminary testing today and are ready to continue shipping co-packaged AI-optimized documentation in the near future. The patterns it established — a directory with an entry point, structured markdown files, generated-not-maintained documentation, and layered discovery mechanisms (standard agent instructions) — generalize cleanly across ecosystems.

We are at an inflection point. The ecosystem moved from “types are someone else’s problem” to “types ship with the package.” The same shift is beginning for AI-optimized documentation. The question is not whether libraries will ship AI-optimized docs, but how quickly the ecosystem converges on a standard for doing so.

This paper proposes that standard. We invite library authors, package registry maintainers, AI tool vendors, and the developer community to adopt, refine, and propagate it.

Ship the Docs in the Box and remember to Think Inside the Box


References

  1. Design Factory AI Documentation Architecture — Internal technical documentation, 2026. To be made publicly available as part of Design Factory’s open-sourcing [TODO: public repository link when available].
  2. Model Context Protocol (MCP) Specification — Anthropic, 2025.
  3. AGENTS.md Convention — Emerging multi-tool convention for AI coding agent instructions.
  4. DefinitelyTyped — Community-maintained TypeScript type definitions
  5. OWASP LLM Top 10 — Prompt Injection risks in LLM applications
  6. npm Provenance — Supply chain security for npm packages
  7. Sigstore — Software signing and transparency for open source
  8. OpenJS Foundation
  9. Python Packaging Authority (PyPA)
  10. S. Bradner — RFC 2119: Key words for use in RFCs to Indicate Requirement Levels
  11. B. Leiba — RFC 8174: Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words
  12. llms.txt — A proposal to standardize on using an /llms.txt file to provide LLM-friendly website content (Jeremy Howard, 2024)
  13. Context7 — Up-to-date Code Docs For Any Prompt
  14. DeepWiki — Cognition AI: AI-generated documentation wikis for public code repositories
  15. GitHub Copilot — Adding repository custom instructions for GitHub Copilot
  16. Cursor — Cursor Rules: project-level instructions for the Cursor AI code editor

This white paper is released for public discussion. We welcome contributions, critiques, and adoption reports from the software engineering community.