Literature Review
Committee annotation
Use the Hypothesis sidebar to highlight text and leave anchored comments. If the sidebar is collapsed, open the small tab on the right edge of the page.
Abstract. Donald E. Knuth’s literate programming paradigm, introduced via his WEB system, envisioned programs as human-readable literature rather than machine-centric code[1]. Despite its conceptual appeal, classical implementations faced persistent adoption barriers[2]. This literature review synthesizes peer-reviewed research, technical reports, and tool documentation to trace the genealogy of paradigms from language-specific ports and independent tools to elucidative programming, formal methods, literate computing, and AI-enhanced systems.
The analysis reveals why traditional literate programming struggled with tooling complexity[2], [3], in contrast to modern notebooks’ dominance in data science, scientific computing, and education[4], [5]. Empirical evidence highlights ongoing reproducibility challenges[6], [7], [8] clear benefits for comprehension and collaboration[5], and AI integration’s promise of substantial productivity gains tempered by security vulnerabilities[9], [10]. This work uncovers key patterns, gaps, and implications to guide future literate programming advancements.
Table of Contents
- Introduction
- Historical Context and Theoretical Foundations
- Literate Programming Paradigms
- Chronological Tool History
- Effectiveness Studies and Empirical Evidence
- Adoption Patterns and Evolution Drivers
- AI Integration and Future Directions
- Discussion and Synthesis
- Conclusion
- References
1. Introduction
1.1 Background and Motivation
Literate programming represents a paradigm shift in software development, introduced by Donald E. Knuth in 1984 with his seminal paper “Literate Programming” in The Computer Journal[1]. Knuth’s vision challenged the traditional approach to programming by proposing that computer programs should be written in a format suited for human understanding rather than machine compilation or execution[1]. This quote captures his philosophy: “Let us change our traditional attitude to the construction of programs: Instead of imagining that our main task is to instruct a computer what to do, let us concentrate rather on explaining to human beings what we want a computer to do.”[1]
Knuth began designing WEB in 1979 with a prototype called DOC, inspired by Pierre-Arnoul de Marneffe’s concept of Holon Programming[11], and refined it into its current form in 1981[1]. Knuth cast WEB into its present form in 1981 while working in the TeX and digital-typography context[1]. The WEB system was documented in Knuth’s September 1983 Stanford report “The WEB System of Structured Documentation”[12]. This system, formally introduced to the wider academic community in the 1984 paper ‘Literate Programming’[1], implemented this vision through a sophisticated framework that allowed programmers to write programs as essays, interleaving code with extensive documentation in a natural order that reflected human thought processes rather than compiler requirements[1]. The system provided two key operations: tangling and weaving[1].
Over four decades, literate programming has evolved far beyond Knuth’s original conception, branching into multiple paradigms and tool ecosystems that serve diverse communities from systems programmers to data scientists, from formal methods researchers to educators, and from academic researchers to enterprise developers[13], [14], [15]. This evolution raises critical questions about the effectiveness of different approaches, the factors driving adoption or resistance, and the future trajectory of literate programming in an era of artificial intelligence.
1.2 Research Objectives and Methodology
This literature review aims to present the evolution of the original literate programming concept, tracing its various branches and key contributions, identifying implementations, and compiling studies on its effectiveness and drawbacks. Despite the advancements, implementations, and studies touting its benefits, we aim to uncover the barriers preventing its widespread adoption and present an argument for continued work.
We constructed the review from a multi-database search of peer-reviewed papers, technical reports, tool documentation, and manuals. The search covered materials published or released from 1979 through June 2026 and combined keyword searching with citation chaining, AI-assisted query expansion, and manual source verification. Search terms included “literate programming,” “computational notebooks,” “Jupyter,” “R Markdown,” “WEB system,” “CWEB,” “noweb,” “elucidative programming,” “reproducible research,” “documentation generators,” “Javadoc,” “Doxygen,” “AI code generation,” and “natural language outlines.” AI tools were used to broaden terminology, identify adjacent tool families, and surface candidate papers more efficiently; final inclusion decisions were made through manual review of each source’s relevance to literate programming genealogy, tool implementation, empirical evidence, or adoption barriers.
2. Historical Context and Theoretical Foundations
2.1 Knuth’s Original Vision (1984)
Donald Knuth introduced literate programming in 1984[1] as a response to what he perceived as a fundamental misalignment in software development: programs were structured for compiler convenience rather than human understanding[1]. His WEB system, developed initially for writing TeX and METAFONT[1], embodied several key principles:
Program as Essay: Programs should be written as literary works, with prose that explains the programmer’s thought process, design decisions, and implementation strategy. The code should emerge naturally from the explanation rather than the explanation being retrofitted to existing code.
Psychological Ordering: Code should be presented in an order that makes sense to human readers, not necessarily in the order required by compilers. The WEB system allowed programmers to write code in any order, with the TANGLE program reorganizing it for compilation[16],[17].
Documentation Integration: Documentation should be integral to the program, not an afterthought. The WEAVE program generated typeset documentation using TeX, treating the program as a publishable technical document[16], [18].
Macro-Based Composition: Programs could be built from named code chunks (macros) that could be defined, refined, and composed hierarchically, supporting top-down design[19].
Knuth’s original WEB system was specifically designed for Pascal and TeX, reflecting his immediate needs for the TeX projects[18]. The system demonstrated the feasibility of literate programming but also revealed challenges that would shape subsequent developments.
2.2 Early Reception and Critique
The software community’s reception of literate programming was mixed[2]. Early critical assessments, documented by Hamer in 1994, emphasized that literate programming appeared “ornate” and impractical for many programmers without better tooling and integration[2]. The overall sentiment is best described in four categories:
Tooling Complexity: The WEB system required learning new markup syntax, understanding the tangling and weaving processes, and integrating these tools into existing development workflows[2], [15]. Although the functionality of WEB and similar literate programming tools was easy to learn, the “style” of literate programming took time to develop[2].
Language Dependence: WEB’s original implementation was tightly coupled to Pascal. While this was appropriate for Knuth’s specific projects, it was later adapted for use with various other languages[2].
Workflow Disruption: Traditional development workflows centered on editing source code files directly. Literate programming required editing WEB files and running additional processing steps, disrupting established practices. The more structured approach to developing code from design documentation required a shift in mindset that many developers found difficult to adopt in practice[2], [15], [20].
Scalability Questions: Concerns arose regarding the adoption of literate programming in larger software development, particularly concerning tool support and the differences from traditional project structures[2], [21].
These early critiques highlighted specific limitations, prompting further consideration of approaches that preserved core principles of literate programming[2].
2.3 Theoretical Foundations and Cognitive Benefits
Despite practical challenges, literate programming rests on sound theoretical foundations in cognitive psychology and software engineering:
Cognitive Load Theory: By presenting code in psychological order with extensive explanation, literate programs reduce cognitive load for readers who must understand complex systems[1]. The narrative structure provides scaffolding that helps readers build accurate mental models[22], [23], [24].
**Dual Coding Theory: **Combining textual explanation with code examples leverages both verbal and visual-spatial cognitive systems, potentially enhancing comprehension and retention[25]. Additionally, adding diagrams that supplement the text and represent the code further drives comprehension and retention[26], [27].
Documentation-Code Consistency: By integrating the documentation and source code, literate programming reduces, rather than fully eliminates, “documentation drift”. This is a common issue where documentation external to the code, such as detailed design descriptions, becomes outdated as the code evolves. Up-to-date documentation provides readers with unified information without the need to reconcile discrepancies in outdated documentation, reducing cognitive dissonance and cognitive load while enabling more efficient construction of accurate mental models[23], [28].
Reflective Practice: Writing literate programs encourages programmers to reflect on their design decisions and assumptions by articulating their reasoning before writing code. Thinking through the problem thoroughly increases the potential for better designs and fewer defects[2]. Documenting design and implementation choices within the codebase itself ensures a deeper understanding of the overall architecture and the rationale behind the system’s implementation[3].
Literate programming also supports cognitive work before, during, and after implementation. Before code is written, prose can function as a planning medium in which the programmer identifies required features, expected inputs and outputs, data structures, invariants, and decomposition choices. During implementation, the literate document can record insights that emerge from testing and debugging, preserving design rationale that is often otherwise lost. After implementation, the same artifact supports later comprehension by making the author’s reasoning available to maintainers rather than leaving them to infer intent from code alone[2], [3], [23].
These communication benefits extend beyond individual comprehension. In team settings, literate artifacts can accelerate onboarding by giving new contributors a structured explanation of design intent, vocabulary, assumptions, and implementation choices. They can also improve continuity across the software lifecycle by preserving rationale for maintainers who were not present during initial development. However, the literature supports improved explanations and shared mental models, but long-term industrial evidence on lifecycle cost reduction remains limited[2], [3], [29].
These theoretical benefits have been partially validated by empirical studies, although the evidence remains mixed and context-dependent[3]. For example, Oman and Curtis provided quantitative evidence that programs presented in a “book format,” where modules are separated into chapters, are easier to comprehend and maintain than traditionally formatted code, even without full literate tooling[2]. Shum and Cook’s classroom experiment demonstrated that literate programming produced more extensive and consistent documentation, though students reported frustrations while debugging[30]. Hamer provides anecdotal reports of fewer bugs due to increased awareness, discouraging sloppy and complex code, along with reduced maintenance costs. The downsides noted were programmer resistance, training needs, and tool incompatibilities[2]. These findings underscore context-dependent outcomes, as discussed further in Section 6. Hamer’s observations on countering a “hack it together” mentality raise the question of whether literate programming restores design rigor diminished in high-level languages relative to the precision demanded by low-level assembly, though no direct empirical link is established[2]. Nevertheless, the emphasis on explanation over mere instruction is a fundamental shift in programming philosophy, compelling developers to prioritize human understanding alongside machine execution[3], [31].
3. Genealogy of Literate Programming Paradigms
The development of literate programming can be viewed as a progression of related paradigms that grew out of Donald Knuth’s original WEB system. Introduced in 1984 for Pascal and TeX, WEB combined code fragments with narrative documentation and used two transformations: tangling to produce compilable code and weaving to generate human-readable documentation[1], [32]. Knuth’s model introduced the influential idea that programs should be written primarily for human understanding, not just for machine execution.
Although conceptually influential, WEB also exposed practical limitations. The system was tightly coupled to specific programming languages, required additional preprocessing steps, and disrupted common development workflows[2], [15], [33]. Subsequent research and tool development sought to preserve the core philosophy of literate programming while addressing these practical challenges. Over time, several distinct branches of literate programming systems emerged, each adapting the original paradigm to new programming environments, development practices, and research domains.
- Language-specific ports that retained WEB’s architecture while adapting it to new programming languages[2], [34].
- Language-independent literate programming systems that simplified syntax and expanded language support[2].
- Elucidative programming approaches that separated documentation and code while maintaining structured relationships between them[35], [36].
- Formal methods integrations that applied literate programming techniques to specification and proof development[37].
- Statistical and reproducible research frameworks that integrate computation with dynamic documents[16], [38].
- Computational notebook environments supporting interactive execution and narrative analysis[5], [39].
- Editor-integrated and bidirectional literate programming systems[40], [41].
- Emerging AI-enhanced literate programming environments[28].
Each branch represents an extension of Knuth’s original vision, or an adaptation of a particular aspect to improve software development practices or related workflows.
3.1 The WEB Lineage: Language-Specific Ports
The earliest extension of literate programming followed directly from Knuth’s original WEB system. These systems retained the tangling and weaving architecture but adapted it to additional programming languages[1], [12].
CWEB, created by Silvio Levy and Donald Knuth in 1987, extended WEB to support the C programming language and later C++[1], [42]. The system introduced preprocessing features compatible with C’s macro mechanisms while preserving the literate programming workflow[1], [42]. CTANGLE generated compilable C code while CWEAVE produced typeset documentation formatted through TeX[42].
Additional ports emerged to support specialized programming communities. FWEB expanded the WEB concept to support Fortran, C, and C++ in scientific computing environments[2], [18]. APLWEB adapted the paradigm to the symbolic and array-oriented syntax of APL, demonstrating the flexibility of the literate programming approach across diverse language structures[43].
These tools demonstrated that literate programming could be applied to multiple programming languages without fundamentally altering the original architecture[15]. However, they also inherited many of WEB’s limitations, including tight coupling to specific languages and the need for separate preprocessing tools[15].
3.2 Language-Independent Tools
Recognizing that language dependence limited the adoption of WEB-style tools[15], a second branch focused on creating language-independent literate programming environments[34].
Norman Ramsey’s noweb system, developed in the late 1980s and formally documented in 1997, became one of the most influential examples[15], [44], [45]. Noweb simplified the syntax for defining code chunks and removed language-specific assumptions, allowing the same literate document structure to generate programs in multiple languages[20], [46]. The tool used a pipeline architecture that integrated easily with standard build systems[15], [45].
Other systems pursued similar goals. FunnelWeb introduced a powerful macro-processing system and allowed output generation independent of specific typesetting tools[2], [47]. Nuweb emphasized minimal syntax and tight integration with LaTeX documentation workflows[2], [33], [48].
These systems represented an important shift from language-specific implementations toward more flexible literate programming frameworks[49]. By simplifying markup and improving compatibility with existing development tools, they reduced the learning curve associated with classical literate programming[41].
3.3 Elucidative Programming Branch
A different direction was proposed by Kurt Nørmark through the concept of elucidative programming in 2000[36]. Rather than embedding code directly within documentation files, elucidative programming separates explanatory documentation from source code while maintaining structured links between them[36], [50].
In this approach, source code remains in conventional files compatible with standard compilers and development tools. Documentation exists as a parallel explanatory structure containing hyperlinks and references that connect narrative explanations to specific code elements. These relationships allow documentation to be regenerated or navigated as the code evolves[35], [51].
This approach eliminates many workflow disruptions associated with tangling and weaving processes, particularly in large projects where related content could be spread across multiple files[35]. Developers can continue to work within familiar development environments while still maintaining extensive explanatory documentation. The approach also enables additional semantic analyses, such as consistency checks between code elements and documentation references.
3.4 Literate Programming and Formal Methods
Another branch emerged through the integration of literate programming with formal methods and proof systems[37]. Formal specifications and proofs often produce artifacts that are difficult for humans to interpret without extensive explanation, making them well suited for literate documentation[52].
Anderson’s 2001 concept of Formal Literate Programming (FLP) proposed integrating formal transformations, specifications, and proofs with narrative explanations within a unified framework[37]. The architecture separated three conceptual layers: a formal layer containing machine-verifiable artifacts, a semi-formal design layer, and an informal explanatory layer written in natural language.
Several proof assistant ecosystems independently adopted literate programming capabilities. The Glasgow Haskell Compiler introduced support for literate Haskell files, allowing code and documentation to coexist in the same document using either “Bird style” or LaTeX-based formatting conventions[53], [54]. Similar approaches were adopted in other proof environments, including Agda, Coq, and Isabelle, where literate modes allow formal proofs to be written as executable research papers[55], [56], [57].
Literate programming concepts have also been explored in the development of safety-critical systems, where traceability, explainability, and rigorous documentation are essential elements of certification. Formal methods are widely recommended or required for high-assurance verification in these domains. For example, IEC 61508-3:2010 recommends formal methods for Safety Integrity Level (SIL) 3 and SIL 4 software systems, while related standards such as DO-178C/DO-333 in avionics, ISO 26262 in automotive systems, and EN 50128 in railway control systems emphasize rigorous traceability between requirements, specifications, verification artifacts, and implementation[58], [59]. The artifacts produced by formal methods (mathematical specifications, model checking results, and theorem-proving scripts) can be difficult for engineers, auditors, and certification authorities to interpret without substantial explanatory context[60], [61] Literate programming provides a natural framework for integrating these formal artifacts with narrative explanations describing design assumptions, refinement steps, and verification rationale[37], [41] Moore and Payne demonstrated this approach by applying WEB-style literate structures to the documentation of assurance arguments for trusted systems, organizing formal specifications, proofs, and explanatory commentary within a unified document structure[52]. Such approaches can improve traceability between requirements and verification artifacts while supporting certification review and independent safety assessment.
These systems demonstrate the usefulness of literate programming for documenting complex formal reasoning processes, particularly in domains where traceability and explanation are essential.
3.5 Statistical Computing Branch
A major branch of literate programming evolved within the statistical computing community, where researchers required tools that integrated narrative explanations, data-analysis code, and generated outputs into reproducible research documents. These tools aligned closely with the broader reproducible research movement in computational science, which emphasizes transparent analytical workflows and executable research artifacts[16].
Friedrich Leisch introduced Sweave in 2002, combining R statistical code with LaTeX documents to produce dynamically generated reports[62]. Sweave enabled authors to embed executable code in research papers, allowing figures, tables, and statistical analyses to be automatically regenerated from the underlying code[62].
Yihui Xie later developed knitr as an enhanced successor to Sweave, introducing improved graphics handling, caching mechanisms, and expanded language support[63]. These tools became central components of reproducible research workflows across statistics, data science, and computational research communities[16].
R Markdown further simplified literate data analysis by combining Markdown syntax with knitr’s execution engine, allowing users to produce documents in multiple output formats including HTML, PDF, and presentation slides[64]. Quarto later extended this ecosystem by supporting multiple programming languages and publishing formats within a unified framework for scientific and technical communication[65].
3.6 Literate Computing
Perhaps the most successful and widely known branch of Literate Programming evolved through interactive computational notebook environments. These systems extended literate programming principles by integrating executable code cells with narrative text and visual outputs within a single interactive document. Unlike classical literate programming systems that use preprocessing tools to generate code and documentation, notebook environments provide immediate feedback through interactive execution of code sections, making them well-suited for exploratory workflows in scientific computing and data science.
The earliest example was Mathematica notebooks, introduced by Stephen Wolfram in 1988. These notebooks provided a graphical environment for combining mathematical computation with explanatory text and visualizations[66].
Other commercial mathematical software packages followed similar styles in explanatory computational pages. Maple introduced worksheets, which allowed users to combine symbolic and numeric computations with explanatory text[67]. Similarly, MATLAB integrated literate programming principles into its environment through the Live Editor, which enables the creation of interactive, formatted documents that interweave code, descriptive text, and visual outputs[68].
Fernando Pérez began developing IPython in the early 2000s as an enhanced interactive environment for Python-based scientific computing[69]. Project Jupyter later emerged in 2014 as a generalization of IPython that supported multiple programming languages through a kernel architecture[70].
Jupyter notebooks rapidly became the dominant environment for data science and computational research, allowing researchers to combine code, narrative explanations, and visual outputs within interactive documents[71]. Later systems, such as JupyterLab, expanded this model into full development environments that support multiple notebooks, terminals, and extensible tools[72].
JupyterHub, launched in approximately 2016[73], enabled multi-user deployments for educational institutions and research teams[74], [75], [76].
3.7 Editor-Integrated Literate Programming Systems
A further branch emerged from editor environments that integrated literate programming capabilities directly into development tools. Rather than relying on external preprocessing steps, these systems embed literate programming features within programmable editors.
Org-mode, introduced for the Emacs editor by Carsten Dominik in 2003[77], provided an extensible plain-text format for organizing documents and executable code blocks[78]. Eric Schulte and Dan Davison later extended Org-mode with Babel, enabling execution of multiple programming languages within a single document and allowing data to be passed between code blocks[79]. This approach introduced the concept of active documents, where narrative explanations, executable code, and generated outputs coexist within a dynamically executable document structure.
The Leo editor, developed by Edward K. Ream beginning around 2001, took a different approach by combining outlining with literate programming[31]. Leo supports optional noweb and CWEB markup within an outline-based editing environment, enabling programmers to organize large programs hierarchically while maintaining literate documentation[20]. Though technically a documentation generator rather than a full literate programming system, Docco spawned numerous ports (Pycco for Python, Rocco for Ruby, Shocco for Shell) and influenced how web developers approach code documentation.
CodeChat occupies a middle position in this branch: the CodeChat Editor keeps ordinary source files editable in conventional programming environments while rendering specially formatted comment/doc blocks as formatted documents. This makes CodeChat best understood as a semi-literate, editor-integrated system rather than a classical tangle/weave tool[3], [49].
Recent systems have extended this concept through bidirectional synchronization between literate documents and extracted source code files. Tools such as Entangled allow developers to edit either the literate document or the generated source files while maintaining consistency between them[41]. This approach addresses a longstanding challenge of classical literate programming: developers often modify generated code directly, breaking synchronization with the literate source[1], [41].
3.8 Automated Documentation Tools
Automated documentation generators should be treated as a related but distinct branch rather than as strict literate programming systems. They invert Knuth’s arrangement: instead of embedding code inside a human-ordered narrative, they embed structured documentation inside source code and generate external reference documents from comments. This makes them weaker on psychological ordering but stronger on workflow compatibility. Their importance for this review is practical: Javadoc, Doxygen, POD, RDoc, pydoc, phpDocumentor, HeaderDoc, and similar tools brought code-documentation co-location into mainstream software practice at a scale classical WEB-style systems never achieved.
The most influential early example was Javadoc, created by Sun Microsystems and shipped as part of the Java platform from its first release in 1996[80]. Javadoc parsed Java source code and generated cross-linked HTML API documentation directly from comment blocks. Its tag vocabulary established a de facto convention for structured documentation comments that became prevalent across the Java ecosystem and was later adopted by many tools. Javadoc was designed to be extensible through a doclet architecture, allowing third parties to redirect its parsed output into alternative formats, such as PDF, or to perform static analysis of an API.
Doxygen[81], first released by Dimitri van Heesch in October 1997, extended this model to a broad range of languages, including C, C++, C#, D, Fortran, Java, PHP, and Python. Its earliest version borrowed code from DOC++, a documentation tool developed by Roland Wunderling and Malte Zoeckler at the Zuse Institute Berlin, before being rewritten in C++. Doxygen adopted a tag syntax closely related to Javadoc’s and substantially expanded the range of output formats to include HTML, RTF, PDF, LaTeX, PostScript, and Unix man pages. A distinctive feature was its ability to perform static analysis of a codebase: by building a parse tree, Doxygen could automatically generate inheritance diagrams, collaboration diagrams, and caller and callee graphs, optionally using the Graphviz dot tool for richer visualizations. These capabilities made it especially popular in large C and C++ projects where understanding structural relationships was as valuable as reading prose descriptions.
The same era produced a proliferation of language-specific generators that applied the comment-extraction model to nearly every major programming community. Perl’s POD (Plain Old Documentation) format, integrated through the perldoc tool from around 1994[82], allowed documentation to be embedded directly in Perl source. Frans Slothouber’s ROBODoc, released in 1995, offered a language-agnostic approach to extracting documentation from tagged source-code comments in virtually any language[83]. Apple introduced HeaderDoc around 2000 to document its C, C++, and Objective-C headers[84]. The scripting and web communities followed with phpDocumentor (Joshua Eichorn, 2000)[85], pydoc (contributed to the Python core by Ka-Ping Yee, 2000)[86], and RDoc for Ruby (Dave Thomas, 2001). Python additionally gained Epydoc[87] (Edward Loper, 2002), while the broader documentation landscape expanded with Greg Valure’s language-independent Natural Docs[88], in the .NET world, NDoc (2004),[89] and later Microsoft’s Sandcastle (2006)[90]. Jeremy Ashkenas started Docco around 2013, popularized a side-by-side documentation style, generating prose and code in parallel columns[91]
Although these tools are sometimes excluded from strict definitions of literate programming because they neither reorder code for human readability nor treat narrative as the primary artifact, they represent the most commercially and practically successful realization of literate programming’s central insight: that documentation and source should be maintained together rather than as separate, divergent artifacts. By embedding documentation within the code and automating its extraction, these generators made structured documentation a routine part of mainstream software engineering, achieving an adoption far broader than Knuth’s original WEB system ever attained.
3.9 AI-Enhanced Literate Programming
The most recent branch of literate programming research explores the integration of large language models and AI-assisted programming tools[28]. These systems extend literate programming principles by treating natural language descriptions not only as documentation but also as executable specifications that guide automated code generation.
Shi et al. proposed the concept of natural language outlines for code generation, in which developers structure software systems using hierarchical textual descriptions that guide automated code synthesis[92]. This approach resembles literate programming in that explanatory structure precedes implementation details while allowing automated systems to generate supporting code artifacts. This process also works the opposite way where the LLM produces outlines that can be verified by humans and aid in the code review process..
Modern notebook environments have also incorporated AI-assisted capabilities[93]. Tools such as Jupyter AI and AI-enabled development environments allow users to generate code, explanations, and documentation interactively within literate programming environments.
Although these approaches remain experimental, they suggest a potential convergence between literate programming, interactive development environments, and AI-assisted software engineering. In such systems, natural language explanations, executable code, and automated reasoning tools may coexist within unified development environments.
4. Chronological Tool History
This section provides a complete chronological timeline of literate programming tools from 1981 to 2025, organized by era and including specific dates, creators, and key features.
4.1 Foundational Era (1981-1990)
| Year | Tool | Creator(s) | Affiliation | Language(s) | Key Innovation | Status |
|---|---|---|---|---|---|---|
| 1981-1984[1], [12] | WEB | Donald E. Knuth[1] | Stanford University[1] | Pascal, TeX | Original literate programming system with TANGLE and WEAVE[1] | Historical |
| 1987 | CWEB | Silvio Levy, Donald Knuth[2], [42] | Princeton/Stanford | C, C++[2] | First major language port of WEB[2] | Maintained |
| 1988 | Mathematica Notebooks | Stephen Wolfram[66], [94] | Wolfram Research[66] | Mathematica[95] | First GUI-based computational notebook[94], [95] | Active |
| 1989 | noweb | Norman Ramsey et al.[44], [45], [53] | Princeton University[45] | Language-agnostic[44], [53] | Simplified, language-independent literate programming[44], [53] | Available |
4.2 Language-Independent Era (1990-2002)
| Year | Tool | Creator(s) | Language(s) | Key Innovation | Status |
|---|---|---|---|---|---|
| 1992[47] | FunnelWeb | Ross Williams[47] | Language-agnostic[47] | Improved macro system[2] | Historical |
| 1993 | APLWEB | Christoph von Basum | APL | APL-specific literate programming | Historical |
| 1996 | Leo Editor | Edward K. Ream | Language-agnostic | Outlining-based literate programming | Active |
| ~1997 | FWEB | Unknown | Fortran, C, C++, Ratfor | Multi-language scientific computing | Historical |
| 2000 | nuweb | Briggs, Ramsdell, Mengel | Language-agnostic | LaTeX-oriented, minimal markup | Available |
| 2000 | CLiP | Ammers et al. | Language-agnostic | Style-based extraction | Research |
| 2021 | Codestrate v2 | CAVI at Aarhus University | JavaScript, TypeScript, Python, Ruby, Lua, HTML | Online, collaborative | Archived |
4.3 Elucidative Programming Era (2000-2012)
| Year | Tool/Concept | Creator(s) | Affiliation | Key Innovation | Status |
|---|---|---|---|---|---|
| 2000 | Elucidative Programming | Kurt Nørmark | Aalborg University | Separate documentation/code with mutual navigation | Research |
| 2000 | Elucidative Scheme Environment | Kurt Nørmark | Aalborg University | Prototype elucidative environment | Research |
| 2012 | DEFT | Wilke et al. | Unknown | Model-based elucidative development | Research |
4.4 Statistical Computing Era (2002-2022)
| Year | Tool | Creator(s) | Affiliation | Language(s) | Key Innovation | Status |
|---|---|---|---|---|---|---|
| 2002 | Sweave | Friedrich Leisch | Vienna University of Technology | R, LaTeX | Dynamic statistical reports[62] | Superseded |
| 2011-2012 | knitr | Yihui Xie | Iowa State University | R, Python, others | Comprehensive reproducible research tool[63] | Active |
| 2012-2014 | R Markdown | RStudio team (J.J. Allaire) | RStudio | R, Markdown | Accessible literate programming with Markdown | Active |
| 2016 | Bookdown | Yihui Xie | RStudio | R, Markdown | Long-form document authoring[96] | Active |
| 2017 | Blogdown | Yihui Xie | RStudio | R, Markdown | Website and blog creation[97] | Active |
| 2018 | Distill | RStudio team | RStudio | R, Markdown | Scientific web publication[98] | Active |
| 2022 | Quarto | Posit team (Allaire, Teague) | Posit | R, Python, Julia, JS | Language-agnostic publishing[65] | Active |
4.5 Literate Computing Era (2001-2020)
| Year | Tool | Creator(s) | Affiliation | Language(s) | Key Innovation | Status |
|---|---|---|---|---|---|---|
| 2001 | IPython | Fernando Pérez | University of Colorado Boulder | Python | Enhanced interactive Python shell | Active |
| 2003 | Org-mode | Carsten Dominik | University of Amsterdam | Plain text | Emacs-based outlining and organization | Active |
| 2010-2012 | Babel | Eric Schulte, Dan Davison | University of New Mexico, Counsyl | Multi-language | Active documents with Org-mode | Active |
| 2013 | Apache Zeppelin | Moon Soo Lee | NFLabs | Multi-language | Big data analytics notebooks | Active |
| 2013 | Databricks Notebooks | Databricks founders | Databricks | Scala, Python, R, SQL | Enterprise big data notebooks | Active |
| 2014 | Project Jupyter | Pérez, Granger, et al. | UC Berkeley, Cal Poly | Multi-language | Language-agnostic computational notebooks | Active |
| 2015-present | CodeChat / CodeChat Editor | Bryan A. Jones et al. | Mississippi State University / open-source project | Multi-language source files | Semi-literate programmer’s word processor; renders comment/doc blocks beside ordinary source code[3] | Active |
| 2015 | Hydrogen | nteract | Community | Multi-language | Atom editor Jupyter integration | Limited |
| 2016 | JupyterHub | Jupyter team | Jupyter Project | Multi-language | Multi-user Jupyter server | Active |
| 2016 | Kaggle Kernels | Kaggle (Google) | Kaggle | Python, R | Competition and learning notebooks | Active |
| 2016 | MATLAB Live Scripts | MathWorks | MathWorks | MATLAB | Interactive MATLAB documents | Active |
| 2016 | nteract | Kyle Kelley, Safia Abdalla | Community | Multi-language | Desktop Jupyter application | Active |
| 2016 | Livebook | Adam Wiggins, Orion Henry, Brett Beutell, Lucía Santamaría | Ink & Switch | Python, JavaScript, Go, browser/web stack | IPython-compatible live notebook with live coding, realtime collaboration, and WYSIWYG prose editing | Experimental prototype |
| 2017 | Google Colaboratory | Google Research | Python | Free cloud-based Jupyter with GPU | Active | |
| 2017 | Codestrates v1 | Roman Rädle, Midas Nouwens, Kristian Antonsen, James R. Eagan, Clemens N. Klokmose | Aarhus University; LTCI / Télécom ParisTech / Université Paris-Saclay | HTML, CSS, JavaScript, JSON | Collaborative, reprogrammable literate computing built on Webstrates | Research prototype; superseded by Codestrates v2 |
| 2018 | JupyterLab | Jupyter team | Jupyter Project | Multi-language | Next-generation Jupyter interface | Active |
| 2018 | Observable | Bostock, Ashkenas, MacWright | Observable | JavaScript | Reactive JavaScript notebooks | Active |
| 2019 | Deepnote | Jakub Juhas, Filip Kočica | Deepnote | Python | Collaborative data science notebooks | Active |
| 2019 | Polynote | Netflix | Netflix | Scala, Python, SQL | Multi-language notebooks for data science | Active |
| 2019 | Iodide | Brendan Colloran and Mozilla contributors | Mozilla | JavaScript, Python via Pyodide, CSS, Markdown/HTML | Browser-native scientific reports that bundle a readable presentation with editable code and client-side computation | No longer actively maintained |
| 2020 | VS Code Notebooks | Microsoft | Microsoft | Multi-language | Native notebook support in VS Code | Active |
4.6 Automated Documentation Tools
| Year | Tool | Creator(s) | Languages(s) | Key Innovation | Status |
|---|---|---|---|---|---|
| 1995/1996 | Javadoc | Sun Microsystems; Doug Kramer documents the design rationale | Java | Generated hyperlinked API documentation from structured source comments; normalized tags such as @param, @return, and @throws | Active in Java toolchain |
| 1995 | ROBODoc | Frans Slothouber | Language-agnostic | Extracted documentation from tagged comments across many languages | Available/open source |
| 1997 | Doxygen | Dimitri van Heesch | C, C++, Java, Python, Fortran, PHP, C#, and others | Multi-language documentation generator with cross references, diagrams, and multiple output formats | Active |
| 2000 | HeaderDoc | Apple | C, C++, Objective-C | API documentation from structured comments in Apple headers | Historical/available |
| 2000 | phpDocumentor | Joshua Eichorn and phpDocumentor project | PHP | Javadoc-style documentation conventions for PHP | Active |
| 2000 | pydoc | Ka-Ping Yee; Python standard library | Python | Runtime/source-based documentation for Python modules | Active |
| 2001 | RDoc | Dave Thomas | Ruby | Documentation extraction for Ruby source and README files | Active |
| 2002 | Epydoc | Edward Loper | Python | API documentation extraction for Python packages | Historical |
| 2003 | Natural Docs | Greg Valure | Language-agnostic | Documentation from plain-language comment style across languages | Active/available |
| 2003 | NDoc | .NET open-source community | .NET languages | Javadoc-like generated documentation for .NET assemblies | Historical |
| 2008 | Sandcastle | Microsoft | .NET languages | Microsoft documentation compiler for managed APIs | Historical/archived |
| 2013 | Docco | Jeremy Ashkenas | Language-agnostic | Lightweight Prose/Code side by side static generator | Active |
4.7 Formal Methods Integration (1990s-2013)
| Year | Tool/System | Language/System | Key Feature | Status |
|---|---|---|---|---|
| Mid-1990s | Literate Haskell | Haskell/GHC | Native .lhs file support in GHC | Active[1] |
| 2000s | coqdoc | Coq | Documentation extraction from proofs | Active[99] |
| 2000s | Literate Agda | Agda | LaTeX-based literate proofs | Active[1] |
| 2001 | FLP | Research[37] | Formal transformations with literate documentation | Research |
| 2013 | Literate sources for content dictionaries | OpenMath | Mathematical specification | Research[100] |
4.8 Modern Bidirectional and AI Era (2005-2025)
| Year | Tool/System | Creator(s) | Language/System | Key Innovation | Status |
|---|---|---|---|---|---|
| 2005/2009 | PyLit | Günter Milde | Python/reStructuredText | Bidirectional text-code conversion for semi-literate Python workflows; first PyPI release visible in 2009 | Available |
| 2013 | Literate CoffeeScript | CoffeeScript project | CoffeeScript/Markdown | Language-level .litcoffee mode: Markdown document with indented executable code blocks | Active as CoffeeScript feature |
| 2016 | Lir | Vassilev, Louhimo, Ikonen, Hautaniemi | Language-agnostic/noweb | Tool-agnostic reproducible data analysis using noweb-style chunks, lir-tangle, lir-make, and lir-weave | Open-source research tool |
| 2023 | Entangled | Johan Hidding | Markdown/language-agnostic | Bidirectional Markdown literate programming | Active |
| 2023 | Jupyter AI | Jupyter Project | Multi-language (Jupyter kernels) | LLM integration for notebooks | Active |
| 2023 | GitHub Copilot Notebooks | GitHub/Microsoft | Multi-language (Jupyter/VS Code) | AI code completion in notebooks | Active |
| 2023 | Cursor IDE | Anysphere Inc. | Multi-language | AI-first code editor with notebook features | Active |
| 2023-2024 | Google Colab AI Features | Python | AI code assistance in Colab | Active | |
| 2024 | Natural Language Outlines | Shi et al. | Language-agnostic/multi-language | LLM-based literate programming | Research |
5. Effectiveness Studies and Empirical Evidence
Empirical evaluation of literate programming has historically lagged behind its conceptual influence. While the paradigm first articulated by Donald Knuth has been widely discussed in software engineering literature, relatively few rigorous quantitative studies have examined its effects on program comprehension, maintainability, or developer productivity. Instead, most evidence comes from classroom experiences, controlled experiments on small systems, and practitioner reports. These studies collectively explore whether integrating narrative explanation with source code improves documentation quality, reduces cognitive load during maintenance, and enhances understanding of complex algorithms. The following section reviews the available empirical research and reported experiences assessing the effectiveness of literate programming across educational and professional contexts.
Classical Literate Programming **Effectiveness. **Empirical evidence for classical literate programming is surprisingly limited, consisting primarily of case studies, experience reports, and small-scale evaluations. In a junior-level course, students using the AOPS literate tool produced substantially more comment words and characters than those using Turbo C, demonstrating increased quantity and quality of documentation[30]. Hurst’s 1996 Australasian CSE paper, “Literate programming as an aid to marking student assignments,” described early classroom experiences in which literate submissions reduced marking effort through structured electronic submission and semi-automated testing[101]. Sulír and Nosal’s 2015 FCSS paper, “Sharing developers’ mental models through source code annotations,” conducted a controlled experiment showing that source code annotations improved program comprehension and reduced maintenance time by 34% compared to unannotated code[102]. Dinmore’s 2012 VL/HCC study, “Design and evaluation of a literate spreadsheet,” applied literate principles to spreadsheets in a controlled user study, finding significant performance improvements in formula comprehension and modification, as well as in dependency layers, compared to traditional spreadsheets[103].
Early Critiques and Adoption Barriers. Early evaluations also highlighted why classical literate programming did not achieve broad adoption. Hamer’s “Literate Programming: A Software Engineering Perspective” [2] provided one of the clearest early critiques, observing that literate programming often appeared “ornate” and impractical, that tool complexity created barriers to adoption[15], that workflow disruption discouraged experimentation[52], and that its benefits were most apparent for complex algorithms requiring extensive explanation[2]. He identified the chief obstacles as limited integration with debuggers and IDEs and the cultural perception that literate programming was more academic than practical[2].
Pieterse et al.’s 2004 paper, ‘A Case for Contemporary Literate Programming,’ surveyed the evolution of literate programming environments and argued that contemporary software-development trends made renewed adoption timely. The paper framed empirical validation as future work, proposing an educational study to measure adoption and documentation outcomes[104].
Elucidative Programming Evidence Gap. Empirical evaluation of elucidative programming remains an evidence gap in this review. Nørmark’s conceptual paper describes source files that remain intact while separate explanatory documentation is linked to named program abstractions; it does not report student outcomes[36].
Computational Notebook Evidence. The computational notebook era has generated significantly more empirical research than classical literate programming. Studies of Jupyter notebooks and related environments provide insight into how narrative programming practices function in modern workflows. Kery and colleagues’ CHI paper “The Story in the Notebook”[14] analyzed how data scientists structure notebooks and found that they frequently construct narrative structure through code-cell ordering, markdown, and progressive exploration rather than through the more deliberate, essay-like formats described in classical literate programming. The study also showed that notebooks function simultaneously as exploratory scratchpads, communication artifacts, and partial records of analysis. This is one reason notebook systems have succeeded where older literate programming tools struggled: they align more naturally with interactive and exploratory work. At the same time, the study highlighted persistent issues, especially regarding execution order, statefulness, and versioning of JSON-based notebook files[14].
“Irre”-Reproducibility Studies. Large-scale reproducibility studies provide some of the strongest empirical evidence in the literature on notebook-based literate workflows. Pimentel et al.’s “A Large-Scale Study About Quality and Reproducibility of Jupyter Notebooks” [8] analyzed more than 1 million notebooks on GitHub and found that only about one-quarter executed without errors, while only a small fraction reproduced identical results in clean environments. Samuel et al.’s “Computational Reproducibility of Jupyter Notebooks from Biomedical Publications” [105] provided even more detailed evidence from a publication-linked sample. The study analyzed 27,271 Jupyter notebooks from 2,660 GitHub repositories associated with 3,467 biomedical publications. Of these, 22,578 notebooks were written in Python, 15,817 declared dependencies in standard requirement files, and 10,388 had all dependencies successfully installed. Only 1,203 notebooks executed without runtime errors, and just 879 notebooks, or 5.56% of the installable subset, reproduced results identical to the originals[105]. In effect, only a very small proportion of the total corpus is reproduced perfectly from a clean environment. Across these studies, the most common reproducibility failures included missing or incomplete dependency specifications, version conflicts, hard-coded file paths, missing datasets, execution-order dependencies between cells, and assumptions tied to local machine environments. These findings are important because they show that notebooks support reproducible workflows in principle, but that actual reproducibility depends heavily on disciplined engineering practices rather than the document format alone.
Quality Attributes of Notebook Workflows. More recent research has attempted to characterize what makes notebooks effective. Vassilev et al.’s PLOS ONE case study “Language-agnostic reproducible data analysis using literate programming” [46] demonstrated that the Lir framework improved both reproducibility and analytical understanding in biomedical data workflows by integrating multiple analysis tools within a unified literate programming environment. Building on this line of research, Huang and colleagues’ study “How Scientists Use Jupyter Notebooks: Goals, Quality Attributes, and Opportunities” [106] identified eight quality attributes that researchers value in notebook environments: clarity, reproducibility, explorability, debuggability, reusability, correctness, performance, and collaboration. Their study also documented practical tactics scientists use to achieve these goals, such as organizing notebooks into coherent narrative sections, documenting assumptions, separating exploratory work from polished analyses, modularizing reusable code, and structuring notebooks so that collaborators and future readers can follow the analytical logic. This work is valuable because it reframes notebooks not simply as tools but as communication artifacts whose quality depends on identifiable engineering and documentation practices.
Comprehension and Navigation Support. Additional studies suggest that notebook usability can be improved through tooling that strengthens structure and navigation. The paper “Enhancing Comprehension and Navigation in Jupyter Notebooks with Static Analysis” [107] introduced HeaderGen, a system that automatically generates markdown headers for notebook code cells using a machine-learning operations taxonomy. The reported user study found that HeaderGen helped participants complete comprehension and navigation tasks more quickly and that they broadly perceived it as useful. This result is noteworthy because it suggests that literate programming benefits can be achieved not only through manual authoring discipline but also through automated or semi-automated tooling that improves document structure post hoc. In that sense, the notebook ecosystem may recover some of the explanatory strengths of classical literate programming through intelligent assistance rather than through fully manual narrative construction. It is also worth noting that, prior to the introduction of HeaderGen[107], there was a proliferation of tools similar to literate programming that generated documentation only after the fact.
Literate Programming in Statistical Computing. The statistical computing community represents one of the most successful applications of literate programming principles. Tools such as Sweave, introduced in “Sweave: Dynamic Generation of Statistical Reports Using Literate Data Analysis,” [62] enabled authors to embed R code directly within LaTeX documents, allowing tables, figures, and statistical results to be regenerated automatically from the underlying code. Later tools such as knitr and R Markdown expanded this approach by simplifying the syntax and supporting multiple output formats, including HTML, PDF, and presentations[108]. Educational research such as “R Markdown: Integrating A Reproducible Analysis Tool into Introductory Statistics”** **[109] reported that students using R Markdown developed reproducible analysis habits more naturally because documentation, code, and output were integrated into a single workflow. Students also benefited from immediate feedback when rendering documents, a clearer separation between explanation and code execution, and professional-looking outputs that improved the communication of results[109]. These systems demonstrate how literate programming principles can be successfully adapted for reproducible research and data analysis workflows.
Research Workflow Integration. The broader idea of treating analyses and experiments as executable documents also appears in work such as Singer’s *“A Literate Experimentation Manifesto.” *[110] Singer argued that computational experiments should include both rich human-written descriptions and an executable sequence of commands capable of recreating the computational setup[110]. This perspective aligns closely with the later success of R Markdown, notebook systems, and reproducible analysis workflows. Its importance in this review lies in its situating literate programming within a broader lineage of executable scientific communication, where the goal is not only better code comprehension but also transparent and reproducible computational reasoning.
Collaborative and Post-Literate Programming. Classical literate programming assumes that an author deliberately crafts explanations alongside code. More recent collaborative work explores whether explanatory context can instead emerge from team communication. Park et al.’s “Post-literate Programming: Linking Discussion and Code in Software Development Teams” [29] proposed capturing discussions from tools such as chat, email, and meetings and linking them to relevant code sections. Their prototype suggested that such links could help team members recover design rationale and contextual knowledge that would otherwise be lost[29]. This approach addresses a major weakness of traditional literate programming, namely the effort required to manually write and maintain narrative explanations[29]. At the same time, it gives up some of the coherence and intentional structure of a carefully authored literary document[29]. The importance of this work lies in broadening the definition of explanatory programming artifacts beyond author-written prose to include team-generated communication traces[29].
Synthesis of Effectiveness Evidence. Across the literature, several patterns emerge consistently. First, there is credible evidence that explanatory structure improves comprehension, communication, and educational outcomes, especially for complex algorithms, analysis workflows, and maintenance tasks[111]. Second, classical literate programming struggled not because its explanations lacked value, but because its traditional tools imposed too much friction relative to mainstream development practices[2]. Third, modern notebook and statistical computing environments demonstrate that literate programming principles can succeed when they are embedded in interactive, low-friction, well-integrated workflows[5]. Fourth, reproducibility remains a major unresolved problem: notebooks and dynamic documents can support reproducible practice, but they do not ensure it without disciplined dependency management, execution hygiene, and documentation of assumptions[6], [112]. Fifth, AI-assisted environments may revive literate programming in a new form by making explanation and code generation more tightly coupled[113], [114], but this promise is accompanied by serious concerns about security, code quality, and trustworthiness[115] [116].
Educational Effectiveness Studies. Several additional empirical studies provide evidence on literate programming’s educational effectiveness. Shum and Cook’s 1994 SIGCSE paper “Using Literate Programming to Teach Good Programming Practices” found that literate programming promoted better code organization and documentation habits among students[30]. Bertholf’s 2000 study at Portland State University directly examined comprehension of literate programs by novice and intermediate programmers[117]. Jones and Mohammadi-Aragh conducted studies examining literate programming in electrical engineering courses[49], [118]. Beck and colleagues extended this educational line by examining writing-to-learn, source-code commenting patterns[119], and codebook methods[120] for analyzing students’ explanatory comments.[118] Birkenkrahe’s 2023 study “Teaching Data Science with Literate Programming Tools” demonstrated the effectiveness of Org-mode in modern data science education[121].
Gaps in the Evidence. Despite a growing body of studies, the evidence base remains uneven. There are still very few long-term industrial or commercial studies comparing literate and non-literate development approaches. Randomized controlled trials are rare[122]. Quantitative studies of maintainability, defect reduction, and long-term documentation quality remain limited. There is also little direct empirical evidence on the effectiveness of AI-enhanced literate programming for complex engineering tasks, and relatively little work examining literate approaches in safety-critical, high-assurance, or regulated domains. These gaps indicate that while the philosophy of literate programming remains highly relevant, many of its strongest claims still require broader and more rigorous validation in modern development environments.
6. Adoption Patterns and Evolution Drivers
6.1 Classical Literate Programming: Limited Adoption
Despite Donald Knuth’s influence and the conceptual elegance of literate programming, classical implementations such as WEB, CWEB, and noweb achieved only limited adoption beyond specialized communities. Early uses were concentrated primarily in academic and research contexts, where the integration of explanation and implementation aligned naturally with pedagogical goals and the publication of algorithms. Literate programming proved particularly useful for presenting complex algorithms in academic papers, documenting research code, and supporting instructional materials that combined explanation with executable examples[1], [16]. In these contexts, the narrative structure of literate programs allowed authors to communicate design rationale and algorithmic structure more clearly than traditional source code listings.
Outside academic settings, adoption was comparatively rare. Although some specialized domains experimented with literate programming, including systems programming projects such as TeX and METAFONT themselves, numerical algorithm implementations, and certain cryptographic systems, industrial uptake remained limited. Organizations often viewed literate programming as incompatible with established development workflows, difficult to integrate with emerging software engineering tools, and burdensome for collaborative development teams. As a result, classical literate programming remained largely a niche technique rather than a widely adopted industry practice.
6.1.2 Barriers to Adoption
For this review, adoption barriers are organized into technical, social, and economic considerations. Pieterse et al. specifically discuss integration with contemporary programming environments and identify empirical adoption research as future work.[104]
Technical barriers were among the most significant. Traditional literate programming required additional preprocessing steps, typically tangling and weaving, that disrupted conventional development workflows and complicated debugging[6]. Developers often had to debug generated source code rather than the original literate documents, creating friction when diagnosing errors[1]. Integration with modern development tools was also limited. Early literate programming environments lacked features that developers came to expect from integrated development environments, including syntax highlighting, refactoring support, code completion, direct version control support, and immediate feedback from linters, formatters, and static analyzers[3]. Some traditional literate tools can still work with these utilities after tangling, because the generated source can be passed to ordinary compilers, linters, and formatters. The harder problem is preserving a smooth feedback loop from tool diagnostics on generated code back to the original literate source, especially in collaborative workflows. Furthermore, language-specific tools such as CWEB and FWEB fragmented the ecosystem by supporting only individual programming languages rather than offering a unified multi-language environment[1], [2].
Social factors also played an important role. Literate programming was frequently perceived as an academic or research-oriented methodology rather than a practical industry technique. This perception limited the emergence of industry champions who might have demonstrated its benefits in large-scale software development. Cultural resistance to documentation-heavy workflows further reduced adoption, as many developers prioritized rapid implementation over extensive narrative explanation[2], [104].
Economic considerations compounded these issues. Organizations could potentially incur upfront costs for training developers, adopting new tools, and restructuring development workflows[2]. Because the return on investment was difficult to quantify up front, management and team leaders were unaware of the practices; many teams were reluctant to invest in literate programming. Maintenance overhead and the limited availability of developers experienced with literate programming techniques further discouraged adoption[2].
6.2 Computational Notebooks: Widespread Adoption
In contrast to classical literate programming systems, computational notebooks such as Jupyter and R Markdown achieved widespread adoption across multiple communities[39], [112]. These environments integrate executable code, narrative text, and visual output within interactive documents, enabling users to develop and communicate computational workflows in a single environment[105], [123].
6.2.1 Adoption Patterns
The strongest adoption occurred in data science, where notebooks became the dominant medium for exploratory data analysis and machine learning experimentation. Interactive execution allows analysts to experiment with code incrementally while immediately visualizing results. Integration with widely used libraries such as pandas, scikit-learn, and TensorFlow further accelerated adoption. Cloud-based platforms, including Google Colab, Kaggle, and Databricks extended notebook accessibility by providing hosted environments that eliminated local installation barriers[124].[125]
Scientific computing communities have also extensively adopted notebook environments. Researchers use notebooks to support reproducible research workflows, to publish executable supplements alongside academic papers, to teach computational methods, and to collaborate on shared analyses across institutions[126], [127], [128]. These practices align closely with the goals of reproducible research, where analytical transparency and repeatability are increasingly emphasized.
Educational institutions similarly embraced notebooks as teaching tools. Interactive tutorials and executable textbooks allow students to experiment with code and observe results immediately, providing rapid feedback that improves engagement and learning outcomes. Cloud-hosted notebook platforms further reduce the technical barriers students face when setting up programming environments[75], [129], [130].
Enterprise adoption has also grown steadily. Organizations use notebooks for data analysis, machine learning development, business intelligence reporting, and internal knowledge sharing. Companies including Google, Microsoft, Netflix, and Databricks have incorporated notebook-based workflows into their data engineering and machine learning platforms [125], [131], [132]
6.2.2 Drivers of Adoption
Several factors contributed to the rapid adoption of notebook-based literate programming environments. Technically, notebooks provide interactive execution, allowing users to run code incrementally without requiring separate compilation or preprocessing steps[133]. Immediate feedback supports exploratory workflows and rapid experimentation[106]. Notebook systems also support rich outputs, including plots, tables, and interactive visualizations, which are essential for data-driven research and analysis[134].
Another critical factor is language flexibility. Jupyter’s kernel architecture allows notebooks to support dozens of programming languages through a common interface, enabling multi-language workflows within a single environment. Cloud integration further reduces technological adoption barriers by enabling users to run notebooks in hosted environments without complex local setup[7].
Social factors also contributed to adoption. Large active communities formed around notebook ecosystems, producing extensive tutorials, educational materials, and shared repositories[39]. Adoption by major technology companies further legitimized notebook-based workflows[135]. Network effects amplified this growth: as more users adopted notebooks, the ecosystem of tools, examples, and community resources expanded rapidly[131].
Economic considerations reinforced these trends. Most notebook platforms are open-source and freely available, lowering financial barriers to adoption[39]. Cloud platforms offer free tiers that make computational resources accessible to students and researchers[125]. In addition, notebooks often accelerate exploratory work, enabling faster iteration cycles and reducing the overhead associated with traditional development environments[5], [129].
6.3 Evolution Drivers
The transition from classical literate programming to modern notebook systems reflects several broader shifts in software development practices.
6.3.1 From Batch to Interactive
One of the most significant shifts was the transition from batch-oriented workflows to interactive execution environments[5], [131], [136], [137]. Classical literate programming systems relied on tangling and weaving processes that generated source code and documentation through separate preprocessing steps. In contrast, notebook environments allow code to be executed directly within the document, providing immediate feedback and supporting exploratory workflows. Early systems, such as Mathematica notebooks, demonstrated the potential of interactive computational documents, and later systems, such as IPython and Project Jupyter, extended this model to general-purpose programming environments.
6.3.2 From Publication to Communication
Another major shift involved the role of literate programming artifacts. Knuth’s original vision emphasized producing publishable program documentation. Modern notebook systems instead prioritize communication and collaboration[138]. Notebooks are frequently used to share analyses, demonstrate computational methods, and communicate results to collaborators or stakeholders[123], [134]. Web-based sharing platforms, collaborative editing environments, and cloud-hosted notebooks enable distributed teams to work together more easily than was possible with classical literate programming tools[71], [135].
6.3.3 From Language-Specific to Language-Agnostic
Classical literate programming systems were typically tied to specific programming languages. In contrast, modern notebook systems support many languages through unified interfaces. Jupyter’s kernel architecture, for example, supports dozens of languages through a common framework[73]. This language-agnostic approach expands the potential user base and allows developers to combine multiple languages within a single analytical workflow.
6.3.4 From Standalone to Integrated
Integration with existing software development tools has also been a key driver of adoption. Modern notebook environments integrate with integrated development environments such as RStudio and Visual Studio Code, cloud platforms such as Google Colab and Databricks, and version control systems such as Git[139]. Package ecosystems and automated pipelines further extend notebook capabilities[7], [140]. These integrations reduce workflow disruption and make literate programming practices easier to incorporate into established development processes[3], [141].
6.3.5 From Expert to Accessible
Finally, modern literate programming environments emphasize accessibility. Classical systems often relied on TeX-based markup and complex preprocessing workflows. In contrast, tools such as R Markdown adopt simple Markdown syntax that is easier for new users to learn. Graphical interfaces, sensible defaults, extensive documentation, and large user communities have further reduced the learning curve associated with literate programming practices[142].
6.4 Domain-Specific Adoption Patterns
The adoption of literate programming approaches varies significantly across domains, reflecting the specific needs of different communities.
In statistical computing, literate programming gained strong traction through tools such as Sweave, knitr, and R Markdown. The statistical community’s emphasis on reproducibility, transparent analysis, and dynamic reporting aligned naturally with literate programming principles[143], [144]. Journals increasingly required reproducible analyses[145], [146], encouraging the adoption of tools that could automatically regenerate figures and results from source data.
Data science communities adopted notebook environments rapidly due to the exploratory nature of their work. Data analysis often involves iterative experimentation, visualization, and collaboration, all of which are well supported by notebook-based workflows. Platforms such as Kaggle further reinforced this culture by providing notebook-based environments for competitions and collaborative experimentation[147].
Scientific computing communities also embraced notebooks to support reproducible research and collaborative analysis. Funding agencies and journals increasingly encourage or require reproducible computational methods, leading researchers to adopt tools that integrate code, documentation, and results within a single executable artifact.
Education represents another major adoption domain. Interactive notebooks provide engaging learning environments where students can experiment with code, visualize results, and receive immediate feedback[148]. This approach lowers barriers to entry for programming education and supports more interactive teaching methods.
6.5 Resistance and Limitations
Taken together, these patterns suggest that modern notebook environments did not replace literate programming so much as reinterpret its principles in a form better aligned with contemporary development practices. The core goal of integrating explanation and implementation remains intact. However, notebook systems emphasize interactivity, accessibility, and collaboration rather than the publication-oriented workflows envisioned in Knuth’s original model.
The evolution from classical literate programming to modern notebook environments therefore reflects a broader shift in software development culture from static program documentation toward interactive, collaborative computational narratives. Understanding these adoption patterns helps explain both the historical limitations of classical literate programming and the rapid success of its modern descendants.
6.5.1 Production Code Concerns
Although computational notebooks have achieved widespread adoption in research and data science workflows, significant concerns remain regarding their suitability for production software development. One commonly cited issue is the presence of hidden execution state. Notebook environments allow cells to be executed out of order, meaning that variables or program state may depend on earlier cells that are no longer visible in the current execution sequence. This behavior can produce hidden dependencies that make code difficult to understand, reproduce, or debug. For production systems where reliability and determinism are essential, such a hidden state introduces risks not present in traditional sequential source code files[7].
Testing practices also present challenges. Standard unit testing frameworks are designed around modular code organized into functions, classes, and packages. In contrast, notebook code is typically structured as sequential cells rather than reusable modules, making automated testing more difficult to implement consistently. As a result, integrating notebooks into conventional testing pipelines may require additional tooling or refactoring code into external modules before it can be properly tested[7], [149].
Version control further complicates collaborative notebook development. Jupyter notebooks are stored in JSON format, which captures both code and metadata, including cell outputs and execution state. This representation is not well-suited to traditional version control workflows, where developers rely on line-based diffs and merges to review changes. As a result, comparing notebook revisions or resolving merge conflicts can be substantially more difficult than with plain-text source files, creating collaboration challenges for teams using Git-based development practices[149].
Scalability is another frequently cited limitation. Notebook environments are optimized for exploratory analysis and small to medium-sized workflows rather than large-scale software systems. Production software development typically relies on modular architectures, abstraction layers, and clear separation of concerns. These architectural principles are difficult to enforce within notebook documents, which tend to encourage linear execution and tightly coupled code blocks. Consequently, notebooks often struggle to scale effectively when applied to large codebases or complex software systems[6] [149].
6.5.2 Software Engineering Practices
Beyond structural limitations, several studies suggest that notebook-based workflows may encourage weaker adherence to traditional software engineering practices. Compared with conventional source code modules, notebooks often contain code that places less emphasis on modular design, abstraction, robust error handling, and systematic testing[6], [8]. Despite their association with literate programming, notebooks may also contain limited explanatory documentation beyond minimal markdown descriptions or exploratory notes[6], [94].
Maintenance challenges further compound these issues. Because notebooks rely on interactive execution, program behavior may depend on the order in which cells were run rather than the order in which they appear in the document. Over time, this can create execution-order dependencies and hidden state, making notebooks difficult to maintain or reuse. Refactoring support is also limited compared with modern integrated development environments, making it harder to restructure code as projects evolve. In addition, extracting reusable components from notebooks into maintainable software modules often requires substantial manual effort, particularly when exploratory analysis code has grown organically over time[6], [7].
6.5.3 Cultural Resistance
In addition to technical concerns, notebooks face cultural resistance in some segments of the software development community. Developers working in systems programming domains, for example, often prefer traditional toolchains that emphasize explicit build processes, strongly typed languages, and tightly controlled execution environments[150], [151]. Within these contexts, notebook-based workflows are sometimes viewed as inappropriate for low-level programming tasks or performance-critical systems[41], [152].
Enterprise software development teams may also restrict the use of notebooks in production environments[149]. Concerns about maintainability, testing, version control integration, and long-term code quality have led some organizations to limit notebooks to exploratory analysis and prototyping phases rather than allowing them within production codebases[149].
Open source communities similarly express reservations. Some projects discourage notebook contributions because the JSON-based file format complicates code review, version comparison, and collaborative editing. These challenges can make it more difficult for maintainers to evaluate changes and maintain consistent development practices across contributors[153].
6.6 Synthesis: Success Factors
Comparing the limited adoption of classical literate programming systems and the widespread success of modern computational notebooks reveals several factors that influence the adoption of explanatory programming environments. Although both approaches share the underlying principle of integrating narrative explanation with executable code[13], [39], their adoption trajectories diverged substantially due to differences in workflow compatibility, accessibility, and ecosystem integration[14], [41].
One of the most significant factors is workflow compatibility. Classical literate programming systems such as WEB and CWEB required developers to adopt specialized authoring styles and preprocessing workflows, including tangling and weaving, before code could be compiled or executed[1], [2]. These requirements disrupted established development practices and created friction in debugging and maintenance processes[2]. In contrast, notebook environments integrate directly into interactive development workflows, allowing users to execute code incrementally within the same document where explanations and results are recorded[154]. This alignment with existing exploratory workflows, particularly in data science and scientific computing, significantly reduced the barriers to adoption[152].
Immediate and visible value also plays an important role in technology adoption. Notebook systems provide immediate feedback through interactive execution and visual output, allowing users to see the results of code changes instantly. This rapid feedback loop supports experimentation, debugging, and exploratory analysis, making the benefits of notebook-based workflows apparent from the first use[154]. By contrast, the advantages of classical literate programming, such as improved documentation quality or enhanced program comprehension, often emerged only over longer time horizons and were therefore more difficult for developers and organizations to justify in terms of short-term productivity[2].
Low barriers to entry further distinguish successful literate programming environments. Modern notebook platforms require minimal setup and frequently operate within web browsers or cloud-hosted environments such as Google Colab and Kaggle[125]. Classical literate programming systems, in contrast, often required familiarity with TeX-based markup, specialized tooling, and additional build processes, which increased the learning curve for new users[2], [3]. Simplified syntax in systems such as Markdown-based notebooks and R Markdown significantly lowered this barrier, enabling wider participation from students, researchers, and practitioners without extensive software engineering backgrounds[108], [109].
Tool integration also strongly influences adoption outcomes. Notebook environments integrate with widely used programming languages, libraries, and development environments, including Visual Studio Code, cloud computing platforms, and collaborative data science tools[139], [141]. These integrations allow literate programming practices to coexist with conventional software engineering workflows rather than replacing them. Classical literate programming systems generally lacked comparable integration with modern development environments, limiting their practicality in large collaborative projects[28], [104].
Community and ecosystem support further amplify these technical advantages. Notebook platforms benefit from large, active communities that produce tutorials, open datasets, reusable notebooks, and extensive documentation[71]. These resources reduce onboarding difficulty and encourage experimentation. Network effects reinforce this growth: as more users adopt notebook environments, the ecosystem of extensions, libraries, and educational resources expands, further accelerating adoption[39].
Domain alignment is another critical factor. Literate programming approaches tend to succeed when they match the natural workflows of a particular field. In statistical computing and data science, analytical workflows inherently involve iterative exploration, visualization, and communication of results[5], [155]. Notebook environments align closely with these practices, allowing researchers to document analytical reasoning alongside executable code and visual outputs[71], [156]. In contrast, traditional software engineering workflows often prioritize modular architectures, automated testing, and strict version control practices, which are less naturally supported by notebook-based documents[6], [153].
Finally, successful technologies often allow incremental adoption rather than requiring wholesale methodological change[157]. Notebook environments can be introduced gradually, starting with exploratory data analysis before integrating them into collaborative research workflows or reporting pipelines[158], [159]. Classical literate programming, by contrast, often required developers to adopt a fully literate workflow from the outset, making partial adoption difficult[1], [2].
Taken together, these factors help explain why modern computational notebooks have achieved widespread adoption while classical literate programming systems remained largely confined to specialized domains. The difference lies not in the underlying philosophy of integrating explanation and implementation, but in how effectively that philosophy was adapted to contemporary development practices, tooling ecosystems, and user expectations.
This comparison also highlights an important lesson for future programming environments: theoretical advantages, endorsements by influential figures, or improved documentation quality alone are insufficient to drive widespread adoption. Instead, successful tools must integrate smoothly with existing workflows, provide immediate and visible benefits, and reduce the cognitive and technical barriers practitioners face. Even the prestige of Knuth’s original proposal could not overcome these practical constraints[14], [20].
7. AI Integration and Future Directions
7.1 AI-Enhanced Literate Programming (2023-2025)
The AI literature is relevant to this review only when it changes the relationship between prose, code, explanation, and verification. General AI-assisted programming studies show that coding assistants can affect productivity and code quality, but they do not by themselves constitute literate programming. The more directly relevant work is narrower: systems that use natural language as a durable specification, maintain bidirectional links between explanations and implementation, generate or critique notebook narratives, or verify whether code and prose remain semantically aligned. This section therefore treats AI as a new capability layer on top of literate programming rather than as a separate programming paradigm. Beginning around 2023, commercial and open-source AI coding tools began embedding themselves in notebook environments and code editors, prompting researchers to reconsider literate programming’s core premises[160].
The integration of large language models with literate programming environments represents the newest branch in the field, but the claim should be made carefully. Zhang et al. frame this moment as a “renaissance of literate programming” and propose Interoperable Literate Programming, a prompt/document structure intended to improve LLM-based code generation for larger projects[28]. Shi et al. propose natural-language outlines as a code-adjacent representation that can support understanding, navigation, generation, and bidirectional synchronization between prose and implementation[92]. Sun and Staron’s 2026 Rosetta Code and CodeNet study adds empirical evidence that trillion-parameter models can align natural-language descriptions with code semantics across 1,228 Rosetta Code tasks and 926 languages, though this finding pertains to model capability rather than to production-ready literate environments[161]. Together, these studies suggest that AI may help address long-standing literate programming problems, especially explanation-code alignment and maintenance, while leaving verification, security, and human understanding unresolved concerns.
7.1.1 GitHub Copilot and Notebooks
GitHub Copilot, launched in 2021 and extended to Jupyter notebook environments in 2023, is a commercially prominent AI-literate programming tool. Within notebooks, Copilot provides context-aware code suggestions drawn from preceding cells, generates both code and Markdown documentation cells, implements functions from natural language comments, and supports multiple programming languages through a unified interface. Its design embodies a partial realization of literate programming principles: the AI reads narrative context (comments, Markdown cells, variable names) and produces code aligned with that context, effectively treating the programmer’s prose as a specification.
Early empirical evaluations of AI-assisted coding in enterprise settings provide encouraging but nuanced results. Chatterjee et al.’s ANZ Bank study reported a six-week GitHub Copilot experiment and later large-scale adoption data from about 1,000 engineers; in the controlled challenge tasks, the Copilot group completed tasks 42.36% faster on average, while security results were inconclusive[162]. Bakal et al.’s 2025 study at ZoomInfo, surveying over 400 developers, reported approximately 20% time savings and 72% positive developer satisfaction, while noting security and code-quality concerns that require additional review. These enterprise-scale studies provide the strongest quantitative evidence to date for AI-assisted programming productivity, but they also reveal that the benefits are uneven across task types and that quality assurance remains an open challenge.
7.1.2 Jupyter AI
The official Jupyter AI extension, introduced in 2023[163], integrates LLM capabilities directly into JupyterLab through a chat interface for code assistance, natural language code generation, code explanation, error diagnosis, and support for multiple LLM backends, including OpenAI, Anthropic, and local models. It was first released publicly in 2023. Its design philosophy explicitly prioritizes augmentation over replacement, providing AI assistance while maintaining human control and understanding. This approach aligns closely with literate programming ideals: the programmer remains the author of the narrative, while the AI serves as an intelligent assistant that can generate, explain, or refactor code within the literate document. Jupyter AI represents an institutional endorsement of AI-literate programming integration by the Project Jupyter community itself.
7.1.3 Natural Language Outlines
Perhaps the most theoretically significant development is Shi et al.’s[92] concept of natural language outlines for code. Their paper, “Natural Language Outlines for Code: Literate Programming in the LLM Era,” proposes inverting the traditional literate programming workflow: rather than writing code and then adding explanations, developers write natural language outlines that LLMs use to generate and maintain corresponding code. Changes to either the natural language or the code are bidirectionally synchronized, maintaining consistency between intent and implementation. This approach directly realizes Knuth’s vision of psychological ordering, where developers express their thinking in the order that makes sense to them, while the AI handles the mechanical translation into executable code. The potential benefits include lower barriers to programming (since natural language is more accessible than code), automatic documentation (since explanations are the source artifact rather than an afterthought), and better alignment with human cognitive processes. Rosiene and Rosiene (2024), presenting at IEEE Frontiers in Education, similarly frame AI-aided code generation through a literate programming lens, arguing that the prompt-to-code workflow mirrors the tangle operation in classical literate programming, with the programmer’s intent expressed in English being “tangled” into executable code.[164]
7.2 Security Concerns
The rapid adoption of AI code generation tools has raised significant security concerns that directly affect the viability of AI-enhanced literate programming. Fu et al.’s 2025 empirical study of Copilot-generated code in GitHub projects analyzed 733 code snippets and found that 29.5% of Python and 24.2% of JavaScript snippets contained security weaknesses, including injection vulnerabilities, cryptographic failures, and access control issues[165]. The authors recommend integrating automated security scanning tools with AI code generation workflows. These findings carry important implications for literate programming: if AI-generated code is embedded within explanatory narratives, the documentation may inadvertently lend credibility to insecure implementations, making security flaws harder to detect. Res et al.’s 2025 work on prompt engineering for secure code synthesis found that explicitly requesting secure code in prompts reduced but did not eliminate vulnerabilities, suggesting that technical mitigation alone is insufficient[166]. Together, these studies indicate that AI-enhanced literate programming environments will need to incorporate security verification as a first-class concern, potentially through automated scanning integrated into the literate document workflow.
7.3 Quality and Training Data
The quality of AI-generated code is fundamentally constrained by the quality of the training data. Improta et al.’s 2025 study demonstrates that AI models reproduce patterns from their training corpora, meaning that if the training data contains poor practices, biased patterns, or outdated idioms, these patterns propagate into the generated code[167]. This finding has particular significance for literate programming: the explanatory narratives in literate documents may obscure underlying quality issues if the AI-generated code appears syntactically correct but embodies poor design patterns. Huang et al.’s 2025 quality framework for notebooks identifies practical quality attributes and tactics, and explores how scientists incorporate AI tools into their notebook workflows[106]. Their work has design implications suggesting that effective AI-enhanced literate programming requires explicit quality-feedback mechanisms that help authors evaluate not just whether code runs, but also whether it meets standards of maintainability, efficiency, and correctness.
7.4 Educational Applications
The intersection of AI and literate programming has particular significance for education. Birkenkrahe, in a paper presented at INTED, argues that the rise of AI coding assistants makes literate programming more relevant, not less, to computer and data science education[168]. As AI handles routine code generation, the ability to articulate programming intent, evaluate generated solutions, and maintain coherent explanatory narratives becomes a core professional competency. Pawagi and Kumar’s work on “probeable problems” for beginner-level programming contests explores how to design educational tasks that require genuine understanding and verification even when AI generates initial solutions, ensuring that literate programming skills, reading, explaining, and reasoning about code remain central to learning[169]. These educational perspectives suggest that AI-enhanced literate programming could become the primary modality for teaching programming, shifting the pedagogical focus from syntax mastery to computational thinking, problem decomposition, and clear technical communication.
7.5 Research Directions and Open Challenges
The convergence of AI and literate programming opens several promising research directions. The most technically ambitious is bidirectional AI-enhanced literate programming, in which systems would maintain automatic synchronization between natural language explanations and code, with AI ensuring consistency as either artifact evolves. Such systems would support multi-level explanation generation (from high-level overviews to line-by-line annotations), conversational refinement through dialogue, and verification that code faithfully implements its stated intent. Shi et al.’s natural language outlines represent a first step toward this vision, but significant engineering and research challenges remain, particularly around maintaining semantic consistency across complex, evolving codebases[92].
A second research frontier involves integrating AI with formal methods through a literate programming lens. AI could generate natural-language explanations of formal proofs, synthesize formal specifications from natural-language requirements, assist in proof development with literate documentation, and verify that implementations match their specifications. This direction would address a longstanding gap in the literate programming literature: Knuth’s original vision focused on explaining algorithms to human readers, but never extended to the formal verification of those explanations’ correctness. AI-assisted formal literate programming could bridge this gap, producing documents that are simultaneously readable, executable, and provably correct.
Collaborative and domain-specific AI assistance represent additional directions with practical applications. Notebook environments incorporate team-aware AI assistants that learn organizational conventions, provide consistent code review, extract reusable patterns from shared notebooks, and generate onboarding documentation tailored to new team members’ backgrounds. Domain-specific AI assistance, i.e., AI that understands scientific methods, statistical test assumptions, data analysis best practices, pedagogical objectives, or enterprise compliance requirements, makes literate programming environments more effective across the diverse communities that have adopted notebooks.
Perhaps the most profound emerging direction is prose-first programming, in which developers write natural-language descriptions of their intent and AI systems generate corresponding code. Andrej Karpathy coined the term “vibe coding” in February 2025 to describe a workflow in which programmers interact with AI primarily in natural language, accepting generated code based on whether it produces correct results rather than reviewing it line by line. While Karpathy framed this somewhat provocatively, the underlying paradigm, prose as the primary programming artifact with code as a derived output, represents the most direct realization of Knuth’s original vision yet attempted. Tools such as Claude Code, Cursor, Windsurf, and Aider exemplify this approach, allowing developers to describe desired behavior in prose and iterate through conversation. Birkenkrahe argues that this shift makes literate programming competencies (clear technical writing, articulation of intent, code reasoning) more important than ever, even as the mechanical act of writing code is increasingly automated[168].
These research directions face substantial challenges across technical, social, and methodological dimensions. On the technical side, reliability remains the foremost concern: AI-generated code and explanations cannot yet be fully trusted for critical applications[170], [171], raising questions about how literate documents should convey confidence levels and uncertainty. Maintaining consistency between AI-generated code and prose explanations as both evolve presents a synchronization problem that current tools do not solve. Scalability to large, complex projects with thousands of interdependent modules remains undemonstrated. Integration with existing development workflows, version control systems, and testing frameworks requires significant engineering work that has just recently begun. This concern is reinforced by early empirical evidence from adjacent AI-assisted programming research. While evaluations of tools such as GitHub Copilot report productivity gains and faster completion of routine coding tasks[172], the security and code-quality concerns detailed in Sections 7.2 and 7.3 temper these benefits, a pattern echoed across the broader literature reviewed by Negri-Ribalta et al. [116]. For literate programming, where AI may produce both code and its explanatory rationale, this tension between productivity and reliability suggests that future systems will need verification, review, and traceability mechanisms stronger than those in today’s notebook and code-generation environments.
Social and ethical challenges are equally pressing. The impact of AI assistance on skill development is contested: if novice programmers rely on AI to generate code from prose descriptions, they may never develop the deep understanding of programming constructs that literate programming was originally designed to cultivate. Attribution becomes complex when AI contributes substantially to both code and documentation; questions of intellectual property, academic honesty, and professional responsibility remain largely unresolved. Biases in AI training data propagate into generated code and explanations, potentially embedding systemic inequities in literate documents that carry the imprimatur of human authorship. Privacy concerns arise when proprietary code and sensitive documentation are processed by cloud-based AI services.
The most significant research gap is the near-complete absence of rigorous empirical evaluation of AI-enhanced literate programming as a unified paradigm. While studies have separately examined AI code-generation productivity, notebook reproducibility, and documentation quality, no published work has yet evaluated systems in which AI simultaneously maintains both prose and code in an integrated literate document. Developing best practices for AI-assisted literate programming, when to accept AI suggestions, how to verify generated explanations, how to structure prompts for literate output, represents an open field. Optimal tool designs for AI-enhanced literate programming environments, effective pedagogical approaches for teaching with AI assistance, and comprehensive security frameworks for AI-generated code within literate documents all constitute active research opportunities with substantial practical impact.
8. Discussion and Synthesis
8.1 The Evolutionary Trajectory
The genealogy of literate programming described in Section 3 reveals not a single lineage but a branching tree of related paradigms, each adapting different aspects of Knuth’s original vision to different communities and technical contexts. Nevertheless, a broad trajectory is discernible. The earliest phase, spanning roughly 1984 to 2000, was dominated by systems that took Knuth’s vision most literally: WEB, CWEB, noweb, and FunnelWeb all treated programs as publishable literature, emphasizing beautifully typeset documentation, psychological ordering, and macro-based code composition[15], [49]. These tools found a devoted but small audience among academic programmers and algorithm researchers, and their limited adoption stemmed largely from the workflow disruption and tooling complexity they imposed. Their lasting contribution was to establish the core principles of code-documentation integration, narrative structure, and the idea that programs should be written for human readers, all of which subsequent branches would inherit.
The second phase, roughly 2001 to 2022, saw literate programming’s principles fragment across multiple parallel branches. Statistical computing tools like Sweave, knitr, and R Markdown brought literate practices to data analysts[144]. Interactive notebooks redefined literate computing as an exploratory, conversational medium rather than a publication format[13], [70]. Editor-integrated systems like Org-mode and Babel embedded literate capabilities into existing workflows[79], [173]. Elucidative programming explored keeping code and documentation in separate but linked artifacts[35], [36]. Formal methods communities adopted literate approaches for proof documentation[37], [52]. Each branch sacrificed some aspect of Knuth’s original vision w@hile amplifying others: notebooks traded psychological ordering for interactivity, R Markdown traded generality for domain fit, and elucidative programming traded integration for workflow compatibility. The common thread was a shift from programs-as-literature toward programs-as-communication. Tools were designed not for publication but for exploration, teaching, collaboration, and reproducibility.
The most recent phase, beginning around 2023, introduces AI as a transformative variable[92]. Tools like GitHub Copilot, Jupyter AI, and natural language outline systems add a new dimension to every branch of the literate programming tree[92]. AI can generate code from prose descriptions, produce explanations of existing code, maintain synchronization between documentation and implementation, and lower the barrier to creating literate documents[92]. Whether this constitutes a genuine third era or simply a new capability layer atop the existing branches remains an open question, but it is clear that AI integration is reshaping what literate programming can mean in practice.
8.2 Why Notebooks Succeeded Where Classical Tools Did Not
The adoption factors analyzed in Section 6.6, including workflow compatibility, immediate and visible value, low barriers to entry, tool integration, community support, and domain alignment, together explain the divergent trajectories of classical literate programming and computational notebooks. Rather than restate that analysis, it is worth drawing out the single lesson that unifies it: notebooks succeeded not because they realized Knuth’s vision more faithfully, but because they asked less of their users. Classical systems required developers to change how they worked before receiving any benefit, and the benefits they offered, such as better documentation and improved comprehension, accrued mainly to future readers rather than the author in the moment. Notebooks inverted this bargain, delivering immediate, visible payoff while fitting the exploratory habits data scientists already had.
This reframing matters for interpreting the rest of this review. It suggests that the persistence of literate programming’s principles, even as its original tools faded, reflects not the triumph of the idea on its own merits but its repackaging into forms that lowered adoption cost. The same lens helps anticipate how AI-assisted approaches (Section 7) may or may not succeed: their fate will likely turn less on how well they embody literate ideals than on whether they reduce, rather than add, friction to existing workflows.
8.3 Persistent Challenges
Despite notebooks’ success, several challenges have proven stubbornly persistent. Reproducibility remains the most well-documented problem: as the large-scale studies detailed in Section 5 show, only a small fraction of public notebooks reproduce identical results in clean environments. The root causes are both technical and cultural, including incomplete dependency specifications[6], hidden state from out-of-order cell execution[7], missing data files, hardcoded paths, and environment-specific configurations. But the deeper issue is that notebook culture prioritizes exploration over reproducibility, users treat notebooks as scratch pads rather than archival documents[39], and the tools do not strongly encourage reproducible practices. Technology alone cannot solve this; reproducibility requires a combination of better tooling, cultural norms, institutional incentives, and education.
A related tension exists between accessibility and quality. The low barriers that make notebooks appealing also enable poor practices: global variables[41], duplicated code[94], absent error handling, and ad hoc structure[41]. Standard software engineering practices, including modular design [6], unit testing [153], code review, and version control [149], fit awkwardly with notebook workflows. Jupyter notebooks are stored as JSON files that produce opaque diffs, making collaborative version control difficult[149], [174]. Testing frameworks assume modular source files, not interleaved narrative-and-code cells[6]. Whether these are fundamental limitations of the notebook paradigm or solvable engineering problems remains contested, but they represent significant barriers to notebook adoption in production software environments.
8.4 Theoretical Contributions
This review suggests several theoretical insights. First, literate programming is better understood as a spectrum than as a binary. At one end lies minimally documented code. Moving along the spectrum, one encounters structured comments and docstrings, then integrated narrative and code in classical literate programs, then interactive computational notebooks, and finally AI-enhanced systems where documentation and code may be generated from each other. Different points on this spectrum suit different contexts, audiences, and goals, and no single point is universally optimal.
Second, the term “literate programming” itself encompasses multiple independent dimensions: narrative structure (from linear to psychologically ordered), explanation depth (from minimal to comprehensive), intended audience (self, team, public, students), interactivity (batch versus interactive), and formality (informal notes versus mathematical rigor). Tools and practices can vary along these dimensions independently, which explains why systems as different as WEB and Jupyter can both claim descent from literate programming. Third, the documentation-code relationship takes fundamentally different forms across paradigms: containment (WEB embeds code within documentation), separation (elucidative programming links separate artifacts), and interleaving (notebooks alternate code and text cells). Each model entails different trade-offs for maintenance, tool integration, and workflow compatibility.
8.5 Practical Implications
For tool developers, the historical record points to clear design principles: minimize workflow disruption, provide immediate visible value, keep barriers to entry low, integrate with existing ecosystems, support incremental adoption, and invest in community building. The tools that failed, despite elegant design and powerful capabilities, were those that required rewiring workflows, deferred benefits to the long term, imposed steep learning curves, or existed as standalone systems disparate from the broader development ecosystem.
For practitioners, the choice of literate programming approach should be guided by context. Notebooks excel at exploratory data analysis, research prototyping, teaching, algorithm explanation, and reproducible reporting. They are less suited to production systems requiring high reliability, large codebases demanding modular architecture, or team projects with strict version control and code review requirements. In all cases, best practices include maintaining clear narrative structure, focusing on reproducibility from the outset, keeping code modular even within notebooks, and documenting not just what the code does but why particular approaches were chosen… the reasoning behind the structure.
For researchers, the most pressing needs are rigorous effectiveness evaluations through randomized controlled trials and large-scale observational studies, longitudinal investigations of how literate documents age and are maintained, comprehensive security assessments of AI-generated code within literate contexts, and pedagogical studies comparing literate and conventional approaches to programming education. For educators, notebooks offer genuine opportunities, including lower barriers, immediate feedback, engaging visualizations, but also risks, particularly if students learn to rely on AI-generated code without developing independent programming competence.
8.6 Convergence and Remaining Divergences
Several convergence trends are visible across the branches of literate programming. Language agnosticism has become the norm: from WEB’s Pascal-specific design, tools have moved toward language-independent architectures, with Jupyter supporting over 100 languages through its kernel system[73]. Interactive execution has replaced batch processing across nearly all modern tools[6]. Web-based presentation has supplanted printed output[2]. Cloud platforms like Google Colab and Databricks have made literate environments accessible without local installation[135]. And AI assistance is being added across platforms, from Jupyter AI to Copilot integration in VS Code notebooks[93], [106].
At the same time, significant divergences persist and may prove permanent. Execution models vary: reactive (Observable), manual cell-by-cell (Jupyter), and batch (R Markdown), reflecting genuine differences in how users think about computation. Artifact structures differ between single-file notebooks and multi-file elucidative approaches. Target audiences range from individual researchers to enterprise development teams. And the fundamental tension between accessibility and production-quality engineering shows no signs of resolving into a single solution. These divergences likely reflect real differences in use cases and user needs rather than temporary fragmentation that will converge over time. The future of literate programming is probably not a single tool or paradigm but a diverse ecosystem of approaches, each suited to particular communities, workflows, and goals.
9. Conclusion
9.1 Summary of Findings
This comprehensive literature review traces the multifaceted evolution of literate programming, from Donald Knuth’s pioneering 1984 WEB system to the AI-enhanced computational notebooks defining 2025. Rather than a straight path, it splintered into vibrant branches: language-specific ports, independent tools, elucidative programming[36], formal methods integration, statistical computing[144], [175] interactive notebooks[67], [133], the Emacs ecosystem[1], [79], and cutting-edge AI systems[92], each adapting to unique needs and eras.
Classical literate programming faltered with limited adoption, hampered by workflow upheavals, tool complexity, and steep entry barriers. Computational notebooks, by contrast, exploded in popularity by embedding seamlessly into daily practices, delivering instant value, easing access, and nurturing robust communities[7].
Empirical evidence paints a mixed picture: notebooks boost comprehension, learning, and collaboration[5], [176], while automated documentation eases navigation[51], [177] while a controlled programming task reported 55.8% faster completion[172] and field experiments reported 12.92% to 21.83% more pull requests per week at Microsoft and 7.51% to 8.69% at Accentur[178] Yet challenges loom, from dismal reproducibility[6] and security flaws in AI code[116] to production hurdles[6], [149], with gaps in broad productivity benchmarks[179].
Propelling this journey were pivotal shifts from batch to interactive execution, from publication to communication, from language-specific to agnostic tools, from standalone to integrated ecosystems, from expert-only to accessible interfaces, culminating in AI infusion[180], [181]. Persistent hurdles remain: reproducibility shortfalls, quality-accessibility tensions, production qualms, AI security risks, and maintenance burdens. The review’s broader theoretical contribution, developed in Section 8.4, is to recast literate programming as a spectrum of independent dimensions and documentation-code relationships rather than a single fixed practice.
9.2 Future Outlook
AI’s fusion with literate programming heralds transformation, charting divergent horizons. Optimistically, it demolishes barriers, empowering masses with quality-assured code via automated checks, natural-language creation for novices, and enriched explanations[92]. Pessimistically, it breeds vulnerabilities, dependency, and shallow facades. Realistically, and most likely, AI excels in drudgery, demanding human savvy for thorny feats, spurring verification advances[7] as literate paradigms proliferate across domains with honed niches[28].
9.3 Future Research Agenda
Future research should move from advocacy to evidence. The field needs controlled studies comparing literate and conventional workflows, larger observational studies across different programmer populations, and longitudinal research on how literate artifacts age during maintenance. For notebook-based workflows, research should focus on practical software-engineering support: testing, refactoring, dependency capture, version control, duplicate-code detection, and durable archival formats[7]. For AI-enhanced systems, the central questions are whether natural-language and code representations can remain synchronized over time, whether generated explanations can be verified against implementation behavior, and how security review should be integrated when AI contributes code[92], [116]. Education studies should examine whether AI-assisted literate programming[182] improves explanation[183], debugging[184], and design reasoning[185], [186], or whether it encourages superficial reliance on generated code.
9.4 Open Questions: Knuth’s Unrealized Vision and Future Research
A careful assessment reveals that several core aspects of his original vision remain unrealized, presenting fertile ground for future research. Knuth envisioned programs as literary works with narrative flow, psychological ordering, and integrated documentation that would make programs as readable as well-written essays. While computational notebooks have achieved broad adoption, they realize only fragments of this vision. This section maps the specific gaps between Knuth’s aspirations and current practice, identifying open questions that represent significant research opportunities.
First, psychological ordering remains largely unsupported in modern tools. Knuth’s WEB system allowed programmers to present code in whatever order made conceptual sense to human readers, with the tangle operation handling resequencing for the compiler[1], [12]. Today’s notebooks enforce a linear, top-to-bottom execution model where cell order determines both narrative and computational sequence[6], [134]. This conflation of presentation order with execution order is a fundamental departure from Knuth’s vision. No mainstream tool currently supports the free reordering of code for narrative purposes while maintaining correct execution semantics. Research question: Can AI-assisted systems restore psychological ordering by automatically managing the mapping between human-readable narrative order and machine-executable dependency order?
Second, the quality of explanation in modern literate documents falls far short of Knuth’s standard. Knuth’s own literate programs (TeX, METAFONT) contained carefully crafted prose explaining not just what the code does, but why design decisions were made, what alternatives were considered, and how the code connects to broader algorithmic theory. Contemporary notebooks typically contain terse Markdown headers and minimal inline comments rather than sustained expository prose. The reproducibility shortfalls documented in Section 5, together with Rule et al.’s observation that notebook narratives are often fragmented and disjointed, suggest that the explanatory richness Knuth envisioned is largely absent from current practice. Research question: What tool features, cultural practices, or AI-assisted scaffolding could elevate the quality of explanation in computational notebooks toward Knuth’s literary standard?
Third, programs as self-contained, publishable literature remains an unachieved ideal. Knuth envisioned literate programs that could be read cover-to-cover like books, with a complete narrative arc. Modern notebooks, by contrast, tend to be ephemeral, exploratory artifacts rather than polished publications. While Quarto, R Markdown, and nbdev have made progress toward publishable computational documents, the norm remains informal, working notebooks that accumulate code cells during an analysis session without retrospective narrative structuring. Research question: How can tools encourage the transition from exploratory notebook to polished literate document, and can AI assist in retrospectively adding narrative structure to existing notebooks?
Fourth, the verification of documentation-code consistency has never been adequately solved. In Knuth’s original system, documentation and code lived in the same source file and were processed together, providing a structural guarantee of co-location but not semantic consistency. As programs evolve, explanations can become stale or misleading, a problem that Knuth acknowledged but did not solve. Neither classical nor modern literate programming tools verify that prose explanations accurately describe the code they accompany. Research question: Can AI systems detect semantic inconsistencies between documentation and code, and can bidirectional synchronization (as proposed by Shi et al.) maintain explanation-code alignment as programs evolve?
Fifth, literate programming for large-scale software systems remains virtually unexplored. Knuth demonstrated literate programming on programs of moderate scale. Virtually no published work demonstrates literate programming practices applied to codebases of hundreds of thousands or millions of lines, the scale at which most professional software engineering occurs[28]. Zhang et al.’s (2024) work on LLM-based code generation for large-scale projects provides initial evidence that AI-assisted literate approaches may scale, but this remains largely unvalidated. Research question: What architectural patterns, tool designs, and AI capabilities are needed to make literate programming viable for large-scale, multi-team software projects?
Sixth, the relationship between literate programming and the potential for reduced software maintenance costs[2]. Most literate programming research focuses on the initial creation of programs, but software systems spend the majority of their lifetime in maintenance and evolution[30]. Whether literate documents improve long-term maintainability, facilitate knowledge transfer between developers, or reduce technical debt remains an empirically open question. Research has indicated reduced maintenance costs with literate programs, though the nature of these studies may warrant further investigation[2]. Research question: Does literate programming reduce long-term maintenance costs, and how do AI-enhanced literate documents age compared to traditional code with conventional documentation?
Finally, the emerging prose-first paradigm raises fundamental questions about the future nature of programming itself. If AI can reliably generate code from natural language specifications, the programmer’s role shifts from code author to prose author and AI supervisor. This transformation has profound implications for computer science education, professional identity, intellectual property, and the epistemology of programming knowledge. Whether this shift ultimately fulfills or subverts Knuth’s vision, whether having humans write prose and machines write code constitutes “literate programming” or something categorically different, is perhaps the deepest open question in the field. Research question: What theoretical framework can distinguish AI-assisted literate programming (where humans and AI collaboratively produce integrated prose-and-code documents) from mere AI code generation (where prose serves only as a prompt, not as a durable explanatory artifact)?
9.5 Final Reflections
Forty years after Knuth’s introduction of literate programming, his core insight remains valid: programs should be written for human understanding, not just machine execution[1]. However, the path to realizing this vision has been neither linear nor simple.
Classical literate programming, despite sound theoretical foundations and expert endorsement[1], failed to achieve widespread adoption due to practical barriers[2]. Computational notebooks succeeded by reimagining literate programming for interactive exploration rather than publication, aligning with actual developer workflows and needs[5], [39].
The current AI integration represents another reimagining: rather than programmers writing both code and explanations, AI may assist with both, potentially lowering barriers[116] while raising new concerns about security, quality, and skill development[187].
The future of literate programming likely involves multiple coexisting approaches serving different needs: notebooks for exploration and communication, traditional tools for production systems, AI assistance for routine tasks, and human expertise for complex problems. The key is matching tools and practices to contexts, audiences, and goals rather than seeking a single universal solution.
This field offers rich opportunities: rigorous evaluation of effectiveness, security and quality frameworks for AI-generated code, optimal tool design, educational effectiveness studies, and the integration of formal methods all represent important open questions. The convergence of literate programming, computational notebooks, and AI assistance creates a particularly wide area for investigation.
Knuth’s vision of programs as literature[1] continues to shape an open research horizon, not through the specific tools he created, but through the principles he articulated and the diverse implementations those principles have inspired[188].
References
[1] D. E. Knuth, Literate programming, vol. 27, no. 2. Oxford University Press, 1984, pp. 97–111. doi: 10.1093/comjnl/27.2.97.
[2] J. Hamer, “Literate programming: a software engineering perspective,” in Proceedings Software Education Conference (SRIG-ET’94), Nov. 1994, pp. 282–288. doi: 10.1109/sedc.1994.475349.
[3] B. A. Jones, M. Mohammadi-Aragh, A. J. Barton, D. Reese, and H. Pan, “Writing-to-Learn-to-Program: Examining the Need for a New Genre in Programming Pedagogy,” Jul. 2015, doi: 10.18260/p.25114.
[4] N. Cheimarios, “Scientific software development in the AI era: reproducibility, MLOps, and applications in soft matter physics,” Frontiers in Physics, vol. 13, Nov. 2025, doi: 10.3389/fphy.2025.1711356.
[5] A. Y. Wang, “Interactive programming interfaces for data science collaboration and learning,” Deep Blue (University of Michigan), Jan. 2023, doi: 10.7302/8520.
[6] J. F. Pimentel, L. Murta, V. Braganholo, and J. Freire, “Understanding and improving the quality and reproducibility of Jupyter notebooks,” Empirical Software Engineering, vol. 26, no. 4, pp. 65–65, May 2021, doi: 10.1007/s10664-021-09961-9.
[7] Md. S. Siddik, H. Li, and C. Bezemer, “A systematic literature review of software engineering research on jupyter notebook,” arXiv (Cornell University), pp. 112758–112758, Apr. 2025, doi: 10.48550/arxiv.2504.16180.
[8] J. F. Pimentel, L. Murta, V. Braganholo, and J. Freire, “A large-scale study about quality and reproducibility of jupyter notebooks,” IEEE, 2019, pp. 507–517.
[9] P. Cao, “Jupyter Notebook Attacks Taxonomy: Ransomware, Data Exfiltration, and Security Misconfiguration,” arXiv (Cornell University), Sep. 2024, doi: 10.48550/arxiv.2409.19456.
[10] T. Weber and S. Mayer, “From Computational to Conversational Notebooks,” arXiv (Cornell University), Jun. 2024, doi: 10.48550/arxiv.2406.10636.
[11] P.-A. de Marneffe and D. Ribbens, “Holon programming,” 1973.
[12] D. E. Knuth, “The WEB system of structured documentation,” Sep. 1983, Accessed: Nov. 2025. [Online]. Available: http://attachment-c02.memonic.ch/p/7baa3bec74b522b61c3e6ce44a36bc38/3251dafb-974a-48e2-a722-e4146a263b99/27e5305d5d/cweb.pdf
[13] B. V. Fog and C. N. Klokmose, “Mapping the landscape of literate computing.,” 2019.
[14] M. B. Kery, M. Radensky, M. Arya, B. E. John, and B. A. Myers, “The story in the notebook,” pp. 1–11, Apr. 2018, doi: 10.1145/3173574.3173748.
[15] N. F. Ramsey, “Literate programming simplified,” IEEE Software, vol. 11, no. 5, pp. 97–105, Sep. 1994, doi: 10.1109/52.311070.
[16] R. Gentleman and D. T. Lang, “Statistical Analyses and Reproducible Research,” Journal of Computational and Graphical Statistics, vol. 16, no. 1, pp. 1–23, Feb. 2007, doi: 10.1198/106186007x178663.
[17] H. Pécout, T. Giraud, and S. Rey‐Coyrehourcq, “Le notebook et la programmation lettrée : documenter ses traitements,” HAL (Le Centre pour la Communication Scientifique Directe), May 2022, Accessed: Aug. 2025. [Online]. Available: https://hal.science/hal-03680189
[18] J. A. Krommes, “Fundamental Statistical Descriptions of Plasma Turbulence in Magnetic Fields,” Feb. 2001. doi: 10.2172/775687.
[19] Václav. Rajlich, “Stepwise refinement revisited,” University of Michigan, Jan. 1983. Accessed: Apr. 2025. [Online]. Available: http://hdl.handle.net/2027.42/7207
[20] A. Kacofegitis and N. Churcher, “Theme-based literate programming,” vol. 8, pp. 549–557, Jun. 2003, doi: 10.1109/apsec.2002.1183079.
[21] J. Bentley, D. Knuth, and D. McIlroy, “Programming pearls,” Communications of the ACM, vol. 29, no. 6, pp. 471–483, Jun. 1986, doi: 10.1145/5948.315654.
[22] J. Sweller, “Cognitive load theory,” vol. 55, Elsevier, 2011, pp. 37–76.
[23] S. Chattopadhyay et al., “Make It Make Sense! Understanding and Facilitating Sensemaking in Computational Notebooks,” arXiv (Cornell University), Dec. 2023, doi: 10.48550/arxiv.2312.11431.
[24] J. Wood, A. Kachkaev, and J. Dykes, “Design Exposition with Literate Visualization,” IEEE Transactions on Visualization and Computer Graphics, vol. 25, no. 1, pp. 759–768, Aug. 2018, doi: 10.1109/tvcg.2018.2864836.
[25] A. Paivio, “Intelligence, dual coding theory, and the brain,” Intelligence, vol. 47, pp. 141–158, 2014, doi: 10.1016/j.intell.2014.09.002.
[26] M. Endres, Z. Karas, X. Hu, I. Kovelman, and W. Weimer, “Relating Reading, Visualization, and Coding for New Programmers: A\n Neuroimaging Study,” arXiv (Cornell University), Feb. 2021, doi: 10.48550/arxiv.2102.12376.
[27] M. Petre and A. F. Blackwell, “Mental imagery in program design and visual programming,” International Journal of Human-Computer Studies, vol. 51, no. 1, pp. 7–30, Jul. 1999, doi: 10.1006/ijhc.1999.0267.
[28] W. Zhang et al., “Renaissance of Literate Programming in the Era of LLMs: Enhancing LLM-Based Code Generation in Large-Scale Projects,” arXiv (Cornell University), Dec. 2024, doi: 10.48550/arxiv.2502.17441.
[29] S. Park, A. X. Zhang, and D. R. Karger, “Post-literate programming: Linking discussion and code in software development teams,” pp. 51–53, Oct. 2018, doi: 10.1145/3266037.3266098.
[30] S. Shum and C. Cook, “Using literate programming to teach good programming practices,” 1994, pp. 66–70.
[31] J. D. Palmer and E. Hillenbrand, “Reimagining literate programming,” pp. 1007–1014, Oct. 2009, doi: 10.1145/1639950.1640072.
[32] B. Antunes and D. R. C. Hill, “Reproducibility, Replicability and Repeatability: A survey of reproducible research with a focus on high performance computing,” Computer Science Review, vol. 53, pp. 100655–100655, Jul. 2024, doi: 10.1016/j.cosrev.2024.100655.
[33] P. Briggs, “Nuweb Version 0.87 b: A simple literate programming tool,” Published on the World-Wide Web by preston@ cs. rice. edu, 1992.
[34] J. W. Bruce, B. A. Jones, and M. Mohammadi-Aragh, “A Literate Programming Approach for Hardware Description Language Instruction,” Sep. 2020, doi: 10.18260/1-2—31966.
[35] C. Wilke, A. Bartho, J. Schroeter, S. Karol, and U. Aßmann, “Extended version of elucidative development for model-based documentation and language specification,” Qucosa (Saxon State and University Library Dresden), Feb. 2012, [Online]. Available: http://tud.qucosa.de/id/qucosa\%3A25893
[36] K. Nørmark, “Elucidative programming,” VBN Forskningsportal (Aalborg Universitet), vol. 7, no. 2, pp. 87–105, Jun. 2000, Accessed: Oct. 2025. [Online]. Available: http://vbn.aau.dk/da/publications/elucidative-programming(56fec7c0-8092-11db-8b97-000ea68e967b).html
[37] H. Anderson, “Formalization and ‘literate’ programming,” vol. 10, pp. 39–44, Aug. 2005, doi: 10.1109/apsec.2001.991457.
[38] M. Bannert, “timeseriesdb: Manage and archive time series data in establishment statistics with R and PostgreSQL,” SSRN Electronic Journal, Jan. 2015, doi: 10.2139/ssrn.2617582.
[39] A. Rule, A. Tabard, and J. D. Hollan, “Exploration and explanation in computational notebooks,” pp. 1–12, Apr. 2018, doi: 10.1145/3173574.3173606.
[40] L. Andersen, C. Moy, S. Chang, and M. Felleisen, “Making Hybrid Languages: A Recipe,” arXiv (Cornell University), Mar. 2024, doi: 10.48550/arxiv.2403.01335.
[41] J. Hidding, “Entangled, a bidirectional system for sustainable literate programming,” IEEE, 2023, pp. 1–9. doi: 10.1109/e-science58273.2023.10254816.
[42] D. E. Knuth and S. Levy, “The CWEB System of Structured Documentation (Version 3.64 — February 2002).” Feb. 2002.
[43] L. J. Dickey, “Literate programming in APL and APLWEB,” ACM SIGAPL APL Quote Quad, vol. 23, no. 4, pp. 11–11, Jun. 1993, doi: 10.1145/173834.173841.
[44] A. Johnson and B. Johnson, “Literate programming using noweb,” Linux journal, vol. 1997, no. 42, p. 1, Oct. 1997, [Online]. Available: https://www.cs.tufts.edu/\textasciitilde{}nr/noweb/johnson-lj.pdf
[45] N. F. Ramsey, “The noweb hacker’s guide,” Jan. 1997.
[46] B. Vassilev, R. Louhimo, E. Ikonen, and S. Hautaniemi, “Language-agnostic reproducible data analysis using literate programming,” Plos one, vol. 11, Art. no. 10, 2016, doi: 10.1371/journal.pone.0164023.
[47] R. N. Williams, “FunnelWeb user’s manual,” 1992.
[48] P. Briggs, J. D. Ramsdell, and M. W. Mengel, “Nuweb Version 0.91 A Simple Literate Programming Tool.” 2000. Accessed: Jan. 2026. [Online]. Available: https://sourceforge.net/projects/nuweb/files/older%20nuweb/0.91/
[49] B. A. Jones, J. W. Bruce, and M. Mohammadi-Aragh, “Exploring Literate Programming in Electrical Engineering Courses,” Computers in education journal, vol. 11, no. 2, Jan. 2020, doi: 10.18260/1-1-118.1153-36157.
[50] K. Nørmark, M. R. Andersen, C. N. Christensen, V. Kumar, S. Staun-Pedersen, and K. L. Sørensen, “Elucidative programming in java,” in International Professional Communication Conference, Sep. 2000, pp. 483–495. doi: 10.5555/504800.504871.
[51] A. van Deursen and T. Kuipers, “Building documentation generators,” pp. 40–49, Jan. 1999, doi: 10.1109/icsm.1999.792497.
[52] A. P. Moore and C. N. Payne, “Increasing assurance with literate programming techniques,” IEEE, 1996, pp. 187–198.
[53] T. C. Ruijs, “Towards effective model checking,” 2001. doi: 10.3990/1.9789036515641.
[54] “Part I The Haskell 2010 Language.” Accessed: Jul. 15, 2026. [Online]. Available: https://www.haskell.org/onlinereport/haskell2010/haskellpa1.html#haskellch10.html
[55] P. Castéran, J. Damour, K. Palmskog, C. Pit-Claudel, and T. Zimmermann, “Hydras & Co.: Formalized mathematics in Coq for inspiration and entertainment,” HAL (Le Centre pour la Communication Scientifique Directe), Jun. 2022, [Online]. Available: https://hal.archives-ouvertes.fr/hal-03404668
[56] B. Igried and A. Setzer, “Programming with monadic CSP-style processes in dependent type theory,” vol. 124, pp. 28–38, Aug. 2016, doi: 10.1145/2976022.2976032.
[57] P.-É. Portier and S. Calabretto, “Methodology for the construction of multi-structured documents,” in Balisage series on markup technologies, Aug. 2009. doi: 10.4242/balisagevol3.portier01.
[58] F. Siron, “Methodology for the formal verification of temporal properties for real-time safety-critical applications based on logical time,” HAL (Le Centre pour la Communication Scientifique Directe), Dec. 2023, Accessed: Aug. 2025. [Online]. Available: https://inria.hal.science/tel-04355316
[59] V. Todorov, “Automotive embedded software design using formal methods,” HAL (Le Centre pour la Communication Scientifique Directe), Dec. 2020, Accessed: Sep. 2025. [Online]. Available: https://theses.hal.science/tel-03082647
[60] D. Miller, “Communicating and trusting proofs: The case for foundational proof certificates,” HAL (Le Centre pour la Communication Scientifique Directe), Jan. 2013, Accessed: Sep. 2025. [Online]. Available: https://hal.inria.fr/hal-00772727
[61] A. P. Kaleeswaran, A. Nordmann, T. Vogel, and L. Grunske, “A user study for evaluation of formal verification results and their explanation at Bosch,” Empirical Software Engineering, vol. 28, no. 5, Sep. 2023, doi: 10.1007/s10664-023-10353-4.
[62] F. Leisch, “Sweave: Dynamic generation of statistical reports using literate data analysis,” in COMPSTAT, 2002, pp. 575–580. doi: 10.1007/978-3-642-57489-4_89.
[63] Y. Xie, “knitr: A general-purpose package for dynamic report generation in R.” Jan. 2012. doi: 10.32614/cran.package.knitr.
[64] Y. Xie, J. Allaire, and G. Grolemund, R Markdown: The Definitive Guide. 2018. Accessed: Nov. 2025. [Online]. Available: https://bookdown.org/yihui/rmarkdown/
[65] S. Mati, İ. Civcir, and S. I. Abba, “EviewsR: An R package for dynamic and reproducible research using EViews, R, R markdown and quarto,” The R Journal, vol. 15, no. 2, pp. 169–205, Nov. 2023, doi: 10.32614/rj-2023-045.
[66] S. Wolfram, “Stephen wolfram,” The New Scientist, vol. 192, no. 2578, pp. 51–51, Nov. 2006, doi: 10.1016/s0262-4079(06)61119-6.
[67] S. Park and E. Sekerinski, “A Notebook Format for the Holistic Design of Embedded Systems (Tool Paper),” Electronic Proceedings in Theoretical Computer Science, vol. 284, pp. 85–94, Nov. 2018, doi: 10.4204/eptcs.284.7.
[68] C. Moler and J. I. Little, “A history of MATLAB,” Proceedings of the ACM on Programming Languages, vol. 4, pp. 1–67, Jun. 2020, doi: 10.1145/3386331.
[69] F. Pérez and B. E. Granger, “IPython: a system for interactive scientific computing,” Computing in science & engineering, vol. 9, Art. no. 3, 2007.
[70] M. Ragan-Kelley et al., “The Jupyter/IPython architecture: a unified view of computational research, from interactive exploration to communication and publication.,” AGU Fall Meeting Abstracts, vol. 2014, Dec. 2014, [Online]. Available: https://agu.confex.com/agu/fm14/webprogram/Paper27181.html
[71] A. Y. Wang, A. Mittal, C. Brooks, and S. Oney, “How data scientists use computational notebooks for real-time collaboration,” Proceedings of the ACM on Human-Computer Interaction, vol. 3, pp. 1–30, Nov. 2019, doi: 10.1145/3359141.
[72] D. Rolon-Mérette, M. Ross, T. Rolon-Mérette, and K. Church, “Introduction to anaconda and python: Installation and setup,” The Quantitative Methods for Psychology, vol. 16, no. 5, May 2020, doi: 10.20982/tqmp.16.5.s003.
[73] M. Milligan, “Jupyter as Common Technology Platform for Interactive HPC Services,” Proceedings of the Practice and Experience on Advanced Research Computing, vol. 9143, pp. 1–6, Jul. 2018, doi: 10.1145/3219104.3219162.
[74] O. Freyermuth, K. Kohl, and P. Wienemann, “Unleashing JupyterHub: Exploiting Resources Without Inbound Network Connectivity Using HTCondor,” Computing and Software for Big Science, vol. 5, no. 1, Oct. 2021, doi: 10.1007/s41781-021-00063-1.
[75] J. Stephanie, O. G. Knut A., N. Robert, J. Alice, and B. Stephen, “Jupyter-Enabled Astrophysical Analysis Using Data-Proximate Computing Platforms,” Computing in Science & Engineering, vol. 23, no. 2, pp. 15–25, Feb. 2021, doi: 10.1109/mcse.2021.3057097.
[76] A. Zonca and R. S. Sinkovits, “Deploying Jupyter Notebooks at scale on XSEDE resources for Science Gateways and workshops,” arXiv, 2018, doi: 10.48550/ARXIV.1805.04781.
[77] T. Johnson, “Emacs as a Tool for Modern Science,” Johnson Matthey Technology Review, vol. 66, no. 2, pp. 122–129, Sep. 2021, doi: 10.1595/205651322x16316969040478.
[78] E. Schulte and D. Davison, “Active Documents with Org-Mode,” Computing in Science & Engineering, vol. 13, no. 3, pp. 66–73, Apr. 2011, doi: 10.1109/mcse.2011.41.
[79] V. G. Pinto, “Performance analysis strategies for task-based applications on hybrid platforms,” HAL (Le Centre pour la Communication Scientifique Directe), Oct. 2018, [Online]. Available: https://theses.hal.science/tel-02063804
[80] D. Kramer, “API documentation from source code comments,” pp. 147–153, Oct. 1999, doi: 10.1145/318372.318577.
[81] D. van Heesch, “Doxygen: Documentation Generator.” Accessed: Jun. 09, 2026. [Online]. Available: https://www.doxygen.nl
[82] “perlhist - the Perl history records.” Accessed: Jun. 17, 2026. [Online]. Available: https://perldoc.perl.org/perlhist
[83] F. Slothouber, “ROBODoc: Documentation Extraction Tool.” Accessed: Jun. 09, 2026. [Online]. Available: https://github.com/gumpu/ROBODoc
[84] Apple Inc., “HeaderDoc.” Accessed: Jun. 09, 2026. [Online]. Available: https://github.com/apple-oss-distributions/headerdoc
[85] J. Eichorn, “phpDocumentor.” Accessed: Jun. 09, 2026. [Online]. Available: https://phpdoc.org
[86] K.-P. Yee, “pydoc.” Accessed: Jun. 15, 2025. [Online]. Available: https://docs.python.org/3/library/pydoc.html
[87] E. Loper, “Epydoc.” Accessed: Jun. 15, 2026. [Online]. Available: https://epydoc.sourceforge.net/
[88] G. Valure, “Natural Docs.” Accessed: Jun. 09, 2026. [Online]. Available: https://www.naturaldocs.org
[89] J. Diamond, “NDoc: Code Documentation Generator for .NET.” Accessed: Jun. 15, 2026. [Online]. Available: https://ndoc.sourceforge.net/
[90] Microsoft, “Sandcastle Help File Builder Documentation.” Accessed: Jun. 15, 2026. [Online]. Available: https://ewsoftware.github.io/SHFB/html/bd1ddb51-1c4f-434f-bb1a-ce2135d3a909.htm
[91] J. Ashkenas, “Docco.” Accessed: Apr. 05, 2026. [Online]. Available: https://github.com/jashkenas/docco
[92] K. Shi et al., “Natural language outlines for code: Literate programming in the LLM era,” arXiv (Cornell University), Aug. 2024, doi: 10.48550/arxiv.2408.04820.
[93] A. M. McNutt, C. Wang, R. A. Deline, and S. M. Drucker, “On the design of ai-powered code assistants for notebooks,” 2023, pp. 1–16. doi: 10.1145/3544548.3580940.
[94] A. P. Koenzen, N. A. Ernst, and M. Storey, “Code Duplication and Reuse in Jupyter Notebooks,” arXiv (Cornell University), Jan. 2020, doi: 10.48550/arxiv.2005.13709.
[95] M. A. H. MacCallum, “Computer algebra in gravity research,” Living Reviews in Relativity, vol. 21, no. 1. Springer Science+Business Media, pp. 6–6, Aug. 20, 2018. doi: 10.1007/s41114-018-0015-6.
[96] Y. Xie, “bookdown: Authoring Books and Technical Documents with R Markdown.” Jul. 13, 2016. doi: 10.32614/cran.package.bookdown.
[97] Y. Xie, A. Thomas, and A. P. Hill, blogdown. 2017. doi: 10.1201/9781351108195.
[98] J. Allaire, R. Iannone, A. P. Hill, and Y. Xie, “Distill for R Markdown,” Sep. 2018, Accessed: Nov. 2025. [Online]. Available: https://apreshill.github.io/distill/
[99] S. Matthieu, “An environment for programming with dependent types,” HAL (Le Centre pour la Communication Scientifique Directe), Dec. 2008, Accessed: Aug. 2025. [Online]. Available: https://theses.hal.science/tel-00640052
[100] C. Lange, “Ontologies and languages for representing mathematical knowledge on the Semantic Web,” Semantic Web, vol. 4, no. 2, pp. 119–158, Jan. 2013, doi: 10.3233/sw-2012-0059.
[101] A. J. Hurst, “Literate programming as an aid to marking student assignments,” 1996, pp. 280–286. doi: 10.1145/369585.369650.
[102] M. Sulír and others, “Sharing developers’ mental models through source code annotations,” IEEE, 2015, pp. 997–1006.
[103] M. Dinmore, “Design and evaluation of a literate spreadsheet,” IEEE, 2012, pp. 15–18.
[104] V. Pieterse, D. G. Kourie, and A. Boake, “A case for contemporary literate programming,” 2004, pp. 2–9. doi: 10.5555/1035053.1035054.
[105] S. Samuel and D. Mietchen, “Computational reproducibility of Jupyter notebooks from biomedical publications,” GigaScience, vol. 13, Jan. 2024, doi: 10.1093/gigascience/giad113.
[106] R. Huang, S. Ravi, M. He, B. Tian, S. Lerner, and M. Coblenz, “How scientists use jupyter notebooks: Goals, quality attributes, and opportunities,” pp. 1243–1255, Apr. 2025, doi: 10.1109/icse55347.2025.00232.
[107] A. P. S. Venkatesh, J. Wang, L. Li, and E. Bodden, “Enhancing comprehension and navigation in jupyter notebooks with static analysis,” IEEE, 2023, pp. 391–401. doi: 10.1109/saner56733.2023.00044.
[108] B. S. Baumer and D. Udwin, “R markdown,” Wiley Interdisciplinary Reviews Computational Statistics, vol. 7, no. 3. Wiley, pp. 167–177, Feb. 2015. doi: 10.1002/wics.1348.
[109] B. S. Baumer, M. Çetinkaya-Rundel, A. Bray, L. Loi, and N. J. Horton, “R markdown: Integrating a reproducible analysis tool into introductory statistics,” Technology Innovations in Statistics Education, vol. 8, no. 1, Jan. 2014, doi: 10.5070/t581020118.
[110] J. Singer, “A literate experimentation manifesto,” pp. 91–102, Oct. 2011, doi: 10.1145/2048237.2048249.
[111] C. J. V. Wyk, “Literate programming,” Communications of the ACM, vol. 31, no. 12, pp. 1376–1395, Dec. 1988, doi: 10.1145/53580.315830.
[112] Anonymous, “PLoS Computational Biology,” Feb. 2005, doi: 10.63485/3xe3z-b6552.
[113] F. Dell’Acqua et al., “Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality,” SSRN Electronic Journal, Jan. 2023, doi: 10.2139/ssrn.4573321.
[114] S. Noy and W. Zhang, “Experimental evidence on the productivity effects of generative artificial intelligence,” Science, vol. 381, no. 6654, pp. 187–192, Jul. 2023, doi: 10.1126/science.adh2586.
[115] S. Torka and Ş. Albayrak, “Optimizing AI-Assisted code generation,” arXiv (Cornell University), Dec. 2024, doi: 10.48550/arxiv.2412.10953.
[116] C. Negri-Ribalta, R. Géraud, A. Sergeeva, and G. Lenzini, “A systematic literature review on the impact of AI models on the security of code generation,” Frontiers in Big Data, vol. 7, pp. 1386720–1386720, May 2024, doi: 10.3389/fdata.2024.1386720.
[117] C. Bertholf, “Comprehension of Literate Programs by Novice and Intermediate Programmers,” Jan. 2000. doi: 10.15760/etd.6456.
[118] M. Mohammadi-Aragh, P. Beck, A. J. Barton, and B. A. Jones, “A Case Study of Writing to Learn to Program: Codebook Implementation and Analysis,” Sep. 2020, doi: 10.18260/1-2—31943.
[119] M. Mohammadi-Aragh, P. Beck, A. J. Barton, D. Reese, B. A. Jones, and M. Jankun-Kelly, “Coding the Coders: A Qualitative Investigation of Students’ Commenting Patterns,” Sep. 2020, doi: 10.18260/1-2—30198.
[120] M. Mohammadi-Aragh, P. Beck, A. Barton, D. Reese, B. A. Jones, and M. Jankun-Kelly, “Coding the Coders,” pp. 1078–1078, Feb. 2018, doi: 10.1145/3159450.3162304.
[121] M. Birkenkrahe, “Teaching Data Science with Literate Programming Tools,” Preprints.org, Jul. 2023, doi: 10.20944/preprints202307.1847.v1.
[122] S. Hanenberg and N. Mehlhorn, “Two N-of-1 self-trials on readability differences between anonymous inner classes (AICs) and lambda expressions (LEs) on Java code snippets,” Empirical Software Engineering, vol. 27, no. 2, Dec. 2021, doi: 10.1007/s10664-021-10077-3.
[123] K. Mendez, L. Pritchard, S. N. Reinke, and D. Broadhurst, “Toward collaborative open data science in metabolomics using Jupyter Notebooks and cloud computing,” Metabolomics, vol. 15, no. 10. Springer Science+Business Media, Sep. 14, 2019. doi: 10.1007/s11306-019-1588-0.
[124] V. Potluri, S. Singanamalla, N. Tieanklin, and J. Mankoff, “Notably Inaccessible — Data Driven Understanding of Data Science Notebook (In)Accessibility,” arXiv, 2023, doi: 10.48550/ARXIV.2308.03241.
[125] D. Rosendo, “Methodologies for reproducible analysis of workflows on the edge-to-cloud continuum,” HAL (Le Centre pour la Communication Scientifique Directe), Jun. 2023, [Online]. Available: https://hal.science/tel-04167278
[126] L. Figueiredo, C. Scherer, and J. S. Cabral, “A simple kit to use computational notebooks for more openness, reproducibility, and productivity in research,” PLoS Computational Biology, vol. 18, no. 9, Sep. 2022, doi: 10.1371/journal.pcbi.1010356.
[127] W. Ammar et al., “Construction of the Literature Graph in Semantic Scholar,” Jan. 2018, doi: 10.18653/v1/n18-3011.
[128] Z. Yaniv, B. Lowekamp, H. J. Johnson, and R. Beare, “SimpleITK Image-Analysis Notebooks: a Collaborative Environment for Education and Reproducible Research,” Journal of Digital Imaging, vol. 31, no. 3, pp. 290–303, Nov. 2017, doi: 10.1007/s10278-017-0037-8.
[129] A. Davies, F. Hooley, P. Freeman, I. Eleftheriou, and G. Moulton, “Using interactive digital notebooks for bioscience and informatics education,” PLoS Computational Biology, vol. 16, no. 11, Nov. 2020, doi: 10.1371/journal.pcbi.1008326.
[130] F. Tobar et al., “Data Science for Engineers: A Teaching Ecosystem,” IEEE Signal Processing Magazine, vol. 38, no. 3, pp. 144–153, Apr. 2021, doi: 10.1109/msp.2021.3053551.
[131] B. Carver, J. Zhang, H. Wang, K. Mahadik, and Y. Cheng, “NotebookOS: A Notebook Operating System for Interactive Training with On-Demand GPUs,” arXiv (Cornell University), Mar. 2025, doi: 10.48550/arxiv.2503.20591.
[132] S. A. Russo, S. Bertocco, C. Gheller, and G. Taffoni, “Rosetta: A container-centric science platform for resource-intensive, interactive data analysis,” Astronomy and Computing, vol. 41, pp. 100648–100648, Sep. 2022, doi: 10.1016/j.ascom.2022.100648.
[133] M. Sp, “Litrepl: Literate Paper Processor Promoting Transparency More Than Reproducibility,” arXiv (Cornell University), Jan. 2025, doi: 10.48550/arxiv.2501.10738.
[134] M. Beg et al., “Using jupyter for reproducible scientific workflows,” Computing in Science & Engineering, vol. 23, no. 2, pp. 36–46, Jan. 2021, doi: 10.1109/mcse.2021.3052101.
[135] N. Schaduangrat, S. Lampa, S. Simeon, M. P. Gleeson, O. Spjuth, and C. Nantasenamat, “Towards reproducible computational drug discovery,” Journal of Cheminformatics, vol. 12, no. 1. BioMed Central, pp. 9–9, Jan. 2020. doi: 10.1186/s13321-020-0408-x.
[136] O. Rind, W. Strecker-Kellogg, D. Allan, D. Benjamin, M. Karasawa, and K. Li, “Integrating Interactive Jupyter Notebooks at the BNL SDCC,” EPJ Web of Conferences, vol. 245, pp. 7054–7054, Jan. 2020, doi: 10.1051/epjconf/202024507054.
[137] J. Reppin et al., “Interactive analysis notebooks on DESY batch resources,” Computing and Software for Big Science, vol. 5, no. 1, Jun. 2021, doi: 10.1007/s41781-021-00058-y.
[138] R. Rädle, M. Nouwens, K. Antonsen, J. Eagan, and C. N. Klokmose, “Codestrates : Literate Computing with Webstrates,” Research Portal Denmark, pp. 715–725, Jan. 2017, Accessed: Jul. 2025. [Online]. Available: https://local.forskningsportal.dk/local/dki-cgi/ws/cris-link?src=au&id=au-161c6504-adb2-4352-9545-c731a98c074b&ti=Codestrates%20%3A%20Literate%20Computing%20with%20Webstrates
[139] J. J. Merelo, “Agile (data) science: a (draft) manifesto.,” arXiv (Cornell University), Apr. 2021, doi: 10.48550/arxiv.2104.12545.
[140] A. Mumuni and F. Mumuni, “Automated data processing and feature engineering for deep learning and big data applications: A survey,” Journal of Information and Intelligence, vol. 3, no. 2, pp. 113–153, Jan. 2024, doi: 10.1016/j.jiixd.2024.01.002.
[141] T. Storer, “Bridging the Chasm,” ACM Computing Surveys, vol. 50, no. 4. Association for Computing Machinery, pp. 1–32, Aug. 25, 2017. doi: 10.1145/3084225.
[142] M. J. Kane, X. Jiang, and S. Urbanek, “On the Programmatic Generation of Reproducible Documents,” Journal of Statistical Software, vol. 103, no. 8, Jan. 2022, doi: 10.18637/jss.v103.i08.
[143] J. Ooms, “Possible Directions for Improving Dependency Versioning in R,” The R Journal, vol. 5, no. 1, pp. 197–197, Jan. 2013, doi: 10.32614/rj-2013-019.
[144] V. Orozco et al., “HOW TO MAKE A PIE: REPRODUCIBLE RESEARCH FOR EMPIRICAL ECONOMICS AND ECONOMETRICS,” Journal of Economic Surveys, vol. 34, no. 5, pp. 1134–1169, Sep. 2020, doi: 10.1111/joes.12389.
[145] R. D. Peng, “Reproducible Research in Computational Science,” Science, vol. 334, no. 6060, pp. 1226–1227, Dec. 2011, doi: 10.1126/science.1213847.
[146] M. Fuentes, “Reproducible Research in JASA,” AMSTAT news: the membership magazine of the American Statistical Association, no. 469, p. 17, Jan. 2016, Accessed: Oct. 2025. [Online]. Available: https://dialnet.unirioja.es/servlet/articulo?codigo=6419357
[147] L. Quaranta, F. Calefato, and F. Lanubile, “KGTorrent: A Dataset of Python Jupyter Notebooks from Kaggle,” May 2021, doi: 10.1109/msr52588.2021.00072.
[148] M. Manuguerra and P. Petocz, “Promoting Student Engagement by Integrating New Technology into Tertiary Education: The Role of the iPad,” Asian Social Science, vol. 7, no. 11, Oct. 2011, doi: 10.5539/ass.v7n11p61.
[149] F. Lanubile, F. Calefato, L. Quaranta, M. Amoruso, F. Fumarola, and M. Filannino, “Towards Productizing AI/ML Models: An Industry Perspective from Data Scientists,” pp. 129–132, May 2021, doi: 10.1109/wain52551.2021.00027.
[150] S. Swords, “A Bound-Finding Tool for Arithmetic Terms,” Electronic Proceedings in Theoretical Computer Science, vol. 393, pp. 11–15, Nov. 2023, doi: 10.4204/eptcs.393.3.
[151] D. Frampton et al., “Demystifying magic,” pp. 81–90, Mar. 2009, doi: 10.1145/1508293.1508305.
[152] Y. Wang et al., “Scaling Notebooks as Re-configurable Cloud Workflows,” Data Intelligence, vol. 4, no. 2, pp. 409–425, Jan. 2022, doi: 10.1162/dint_a_00140.
[153] E. Chalstrey, “Developing and Publishing Code for Trusted Research Environments: Best Practices and Ways of Working,” arXiv (Cornell University), Feb. 2022, doi: 10.48550/arxiv.2111.06301.
[154] A. Shome, L. Cruz, D. Spinellis, and A. van Deursen, “Understanding feedback mechanisms in machine learning jupyter notebooks,” arXiv (Cornell University), Jul. 2024, doi: 10.48550/arxiv.2408.00153.
[155] A. McNamara, “Key Attributes of a Modern Statistical Computing Tool,” The American Statistician, vol. 73, no. 4, pp. 375–384, Jun. 2018, doi: 10.1080/00031305.2018.1482784.
[156] D. Ramasamy, C. Sarasua, A. Bacchelli, and A. Bernstein, “Visualising data science workflows to support third-party notebook comprehension: an empirical study,” Empirical Software Engineering, vol. 28, no. 3, Mar. 2023, doi: 10.1007/s10664-023-10289-9.
[157] T. Judd, A. Ryan, E. Flynn, and G. McColl, “If at first you don’t succeed … adoption of iPad marking for high-stakes assessments,” Perspectives on Medical Education, vol. 6, no. 5, pp. 356–361, Aug. 2017, doi: 10.1007/s40037-017-0372-y.
[158] C. Chen et al., “WHATSNEXT: Guidance-enriched Exploratory Data Analysis with Interactive, Low-Code Notebooks,” arXiv (Cornell University), Aug. 2023, doi: 10.48550/arxiv.2308.09802.
[159] X. Li, Y. Zhang, J. Leung, C. P. Sun, and J. Zhao, “EDAssistant: Supporting Exploratory Data Analysis in Computational Notebooks with In-Situ Code Search and Recommendation,” arXiv (Cornell University), Feb. 2022, doi: 10.48550/arxiv.2112.07858.
[160] S. Zhang, J.-P. Wang, G. Dong, J. Sun, Y. Zhang, and G. Pu, “Experimenting a New Programming Practice with LLMs,” arXiv (Cornell University), Jan. 2024, doi: 10.48550/arxiv.2401.01062.
[161] S. Sun and M. Staron, “Literate Programming With LLMs? — A Study on Rosetta Code and CodeNet,” IEEE Transactions on Software Engineering, vol. 52, no. 2, pp. 468–487, Nov. 2025, doi: 10.1109/tse.2025.3629828.
[162] S. Chatterjee, C. L. Liu, G. Rowland, and T. Hogarth, “The impact of AI tool on engineering at ANZ bank an empirical study on GitHub copilot within corporate environment,” Software Engineering, pp. 15–30, Apr. 2024, doi: 10.5121/csit.2024.140702.
[163] “Generative AI in Jupyter,” Jupyter Blog, Aug. 2023, Accessed: Jul. 07, 2026. [Online]. Available: https://blog.jupyter.org/generative-ai-in-jupyter-3f7174824862
[164] J. Rosiene and C. P. Rosiene, “WIP: Focusing on Programmer Literacy in the Time of AI-Aided Code Generation,” pp. 1–4, Oct. 2024, doi: 10.1109/fie61694.2024.10893329.
[165] Y. Fu et al., “Security weaknesses of copilot-generated code in GitHub projects: An empirical study,” ACM Transactions on Software Engineering and Methodology, pp. 1–34, Feb. 2025, doi: 10.1145/3716848.
[166] M. Bruni, F. Gabrielli, M. Ghafari, and M. Kropp, “Benchmarking Prompt Engineering Techniques for Secure Code Generation with GPT Models,” arXiv (Cornell University), Feb. 2025, doi: 10.48550/arxiv.2502.06039.
[167] C. Improta, R. Tufano, P. Liguori, D. Cotroneo, and G. Bavota, “Quality in, quality out: Investigating training data’s role in AI code generation,” arXiv (Cornell University), Mar. 2025, doi: 10.48550/arxiv.2503.11402.
[168] M. Birkenkrahe, “THE ROLE OF AI CODING ASSISTANTS: REVISITING THE NEED FOR LITERATE PROGRAMMING IN COMPUTER AND DATA SCIENCE EDUCATION,” INTED proceedings, vol. 1, pp. 127–132, Mar. 2024, doi: 10.21125/inted.2024.0071.
[169] M. Pawagi and V. Kumar, “Probeable problems for beginner-level programming-with-AI contests,” pp. 166–176, Aug. 2024, doi: 10.1145/3632620.3671108.
[170] S. K. Navneet and J. Chandra, “Rethinking Autonomy: Preventing Failures in AI-Driven Software Engineering,” arXiv (Cornell University), Aug. 2025, doi: 10.48550/arxiv.2508.11824.
[171] D. Cotroneo, C. Improta, and P. Liguori, “Human-Written vs. AI-Generated Code: A Large-Scale Study of Defects, Vulnerabilities, and Complexity,” ArXiv.org, Aug. 2025, doi: 10.48550/arxiv.2508.21634.
[172] S. Peng, E. Kalliamvakou, P. Cihon, and M. Demirer, “The impact of AI on developer productivity: Evidence from GitHub copilot,” arXiv (Cornell University), Jan. 2023, doi: 10.48550/arxiv.2302.06590.
[173] K. Hinsen, “Explaining software and computational methods,” Jun. 2025, doi: 10.59350/g0hyh-qzx40.
[174] S. Hodencq, “Methods and tools for a collaborative and open energy modelling process,” HAL (Le Centre pour la Communication Scientifique Directe), Sep. 2022, [Online]. Available: https://hal.science/tel-03809331
[175] P. Mair, “Thou Shalt Be Reproducible! A Technology Perspective,” Frontiers in Psychology, vol. 7, Jul. 2016, doi: 10.3389/fpsyg.2016.01079.
[176] G. Amoudi and D. Tbaishat, “Interactive notebooks for achieving learning outcomes in a graduate course: a pedagogical approach,” Education and Information Technologies, vol. 28, no. 12, pp. 16669–16704, May 2023, doi: 10.1007/s10639-023-11854-x.
[177] A. Štěpánek, D. Kuťák, B. Kozlíková, and J. Byška, “Helveg: Diagrams for Software Documentation,” IEEE Transactions on Visualization and Computer Graphics, vol. 31, no. 10, pp. 9079–9091, Jul. 2025, doi: 10.1109/tvcg.2025.3589748.
[178] K. Z. Cui, M. Demirer, S. Jaffe, L. Musolff, S. Peng, and T. Salz, “The Productivity Effects of Generative AI: Evidence from a Field Experiment with GitHub Copilot,” Mar. 2024, doi: 10.21428/e4baedd9.3ad85f1c.
[179] Z. Zhou et al., “From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems,” arXiv (Cornell University), Mar. 2025, doi: 10.48550/arxiv.2503.01424.
[180] T. Zhuang and Z. Lin, “The why, what, and how of AI-based coding in scientific research,” arXiv (Cornell University), Oct. 2024, doi: 10.48550/arxiv.2410.02156.
[181] A. Zamfiroiu, C.-B. PASTINICA, C. Crăciun, C. Ciutacu, and T. LUCA, “Exploring Conversational Programming through the CHOP Paradigm and Artificial Intelligence,” Informatica Economica, vol. 29, pp. 5–17, Jun. 2025, doi: 10.24818/issn14531305/29.2.2025.01.
[182] P. Beck and M. Mohammadi-Aragh, “Archimedes: Developing a Model of Cognition and Intelligent Learning System to Support Metacognition in Novice Programmers,” pp. 1–5, Oct. 2020, doi: 10.1109/fie44824.2020.9274133.
[183] P. Beck, M. Mohammadi-Aragh, and C. Archibald, “An Initial Exploration of Machine Learning Techniques to Classify Source Code Comments in Real-time,” Sep. 2020, doi: 10.18260/1-2—32065.
[184] P. Beck and M. Mohammadi-Aragh, “Board 421: Using a Timeline of Programming Events as a Method for Understanding the Introductory Students’ Programming Process,” Feb. 2024, doi: 10.18260/1-2—42758.
[185] P. Beck, M. Mohammadi-Aragh, C. Archibald, B. A. Jones, and A. Barton, “Real-time Metacognition Feedback for Introductory Programming Using Machine Learning,” pp. 1–5, Oct. 2018, doi: 10.1109/fie.2018.8658973.
[186] P. Beck and M. Mohammadi-Aragh, “An Initial Investigation of Design Cohesion as a IDE-based Learning Analytic for Measuring Introductory Programming Metacognition,” Aug. 2024, doi: 10.18260/1-2—46561.
[187] C. P. Chong, Z. Yao, and I. Neamtiu, “Artificial-Intelligence Generated Code Considered Harmful: A Road Map for Secure and High-Quality Code Generation,” arXiv (Cornell University), Sep. 2024, doi: 10.48550/arxiv.2409.19182.
[188] F. Y. Congo, “Proposing a representation model of computational simulations’ execution context for reproducibility purposes,” HAL (Le Centre pour la Communication Scientifique Directe), Dec. 2018, Accessed: Jun. 2025. [Online]. Available: https://tel.archives-ouvertes.fr/tel-02363764