aliases:
- “Harnessing AI Agent Teams with Literate Programming for Structured Software Development: A Comprehensive Literature Review”
Executive Summary
This literature review examines the emerging intersection of AI agent systems, multi-agent coordination, and software development methodologies, with particular attention to the potential integration of literate programming principles. The review synthesizes evidence from 61 studies spanning 2013–2026, covering agentic software engineering, multi-agent architectures, specification-driven development, testing automation, and system integration validation. Key findings reveal that while multi-agent systems demonstrate significant capabilities in autonomous code generation, testing, and deployment—with systems like HyperAgent achieving 33% success rates on industry benchmarks and AutoDev reaching 87.8% test generation accuracy—explicit integration of literate programming principles remains an identified research gap. The review identifies proven approaches to multi-agent coordination (role specialization, formal orchestration substrates, cross-team collaboration), specification-driven workflows, test-driven development with AI agents, and integration validation frameworks. However, critical gaps persist in automated design documentation maintenance, accurate build-time estimation from specifications, and empirical validation of time-to-market improvements. These findings establish a foundation for dissertation research on structured development harnesses that combine agent teams with literate programming elements to achieve predictable build times, accelerated feature velocity, and continuously synchronized design documentation.
Table of Contents
- Introduction
- Background and Theoretical Foundations
- Multi-Agent Team Coordination and Task Breakdown
- Specification-Driven Design with AI Agents
- Testing Requirements and Test-Driven Development with AI Agents
- System Integration Validation
- Structured Development Harness and Workflow
- Automated Design Documentation Maintenance
- Outcomes: Build Time Accuracy, Time-to-Market, and Feature Velocity
- Research Gaps and Future Directions
- Conclusion
1. Introduction
The rapid advancement of large language models (LLMs) and autonomous AI agents has catalyzed a fundamental transformation in software engineering practices. Contemporary research demonstrates that AI agents can autonomously perform complex software development tasks—from requirements analysis and code generation to testing and deployment—with increasing sophistication and reliability (Otoum et al., 2026). This evolution has given rise to what some researchers term “Software Engineering 3.0,” characterized by AI teammates that reshape traditional development workflows (Zhang, 2024). However, as these agentic systems grow more capable, critical questions emerge about how to structure, coordinate, and validate their work to achieve predictable outcomes, maintain design coherence, and ensure that documentation remains synchronized with evolving codebases.
This literature review addresses a specific research gap at the intersection of three domains: (1) multi-agent AI systems for software development, (2) structured development methodologies emphasizing upfront specification and design, and (3) literate programming principles that treat documentation as a first-class artifact interwoven with code. While substantial research exists on agentic coding systems and multi-agent coordination, the explicit integration of literate programming elements—where narrative documentation and executable code form a unified, maintainable whole—remains largely unexplored in the context of AI agent teams.
The dissertation research motivating this review proposes a structured development harness that combines multi-agent coordination with literate programming elements to achieve three key outcomes: (1) more accurate build time estimation through explicit specification and task breakdown, (2) faster time-to-market through coordinated agent teams and automated testing, and (3) accelerated feature velocity through continuously maintained design documentation that reduces cognitive overhead and onboarding friction. This review synthesizes current evidence on each component of this vision, identifies proven approaches and persistent gaps, and establishes a foundation for empirical investigation of integrated literate-agentic development workflows.
The review is organized around eight thematic areas corresponding to the dissertation’s focus: AI agents and literate programming principles, multi-agent team coordination and task breakdown, specification-driven design, testing and test-driven development (TDD) with AI agents, system integration validation, structured development harnesses and workflows, automated design documentation maintenance, and measurable outcomes including build time accuracy, time-to-market, and feature velocity. Each section synthesizes evidence from the top 30 most relevant papers from two comprehensive literature searches, critically evaluates the state of the art, and identifies research gaps that the dissertation aims to address.
2. Background and Theoretical Foundations
2.1 AI Agents in Software Engineering
AI agents in software engineering are autonomous or semi-autonomous systems that perceive their environment, reason about tasks, and take actions to achieve specified goals (Wang et al., 2025). Recent surveys characterize agentic AI for software development as systems capable of planning, tool integration, context management, and iterative refinement without continuous human intervention (Wang et al., 2025). Otoum et al. (2026) conducted a systematic literature review of methods and techniques in agentic software engineering, identifying key capabilities including autonomous code generation, automated testing, requirements engineering, and deployment automation.
The evolution from simple code completion tools to sophisticated agentic systems represents a paradigm shift. Early LLM-based coding assistants provided context-aware suggestions and snippet generation, but contemporary agentic systems demonstrate end-to-end autonomy across the software development lifecycle (Nasir et al., 2025). These systems can interpret natural language requirements, plan implementation strategies, generate code, write tests, debug failures, and iterate toward working solutions with minimal human guidance (Tufano et al., 2024; Rondon et al., 2025).
A critical distinction in the literature separates “vibe coding”—informal, exploratory interaction with LLMs—from “agentic coding,” which emphasizes structured workflows, explicit goals, and systematic validation (Sapkota et al., 2025). Agentic coding systems incorporate planning mechanisms, maintain state across interactions, use tools and APIs to interact with development environments, and employ feedback loops to refine outputs based on test results and execution outcomes (Wang et al., 2025). This structured approach aligns more closely with traditional software engineering rigor while leveraging the generative capabilities of LLMs.
2.2 Literate Programming Principles
Literate programming, introduced by Donald Knuth in 1984, represents a paradigm shift in how software is conceived and documented. The core principle is that programs should be written primarily for human understanding, with executable code extracted as a byproduct of a narrative explanation (Knuth, 1984). In literate programming, the source artifact is a document that interweaves natural language exposition with code fragments, allowing developers to explain the “why” and “how” of their design decisions alongside the implementation itself.
Despite its theoretical appeal, literate programming has seen limited adoption in mainstream software development, often relegated to specialized domains such as scientific computing and reproducible research (e.g., Jupyter notebooks, R Markdown). However, the principles underlying literate programming—treating documentation as primary, maintaining synchronization between explanation and implementation, and structuring code for human comprehension—remain highly relevant to contemporary challenges in software maintainability and knowledge transfer.
In the context of AI agent systems, literate programming principles offer a potential solution to the “black box” problem of generated code. When agents autonomously produce implementations, the rationale behind design decisions, the assumptions embedded in the code, and the relationships between components may not be immediately apparent to human developers who must maintain or extend the system. Integrating literate programming elements into agentic workflows could ensure that generated code is accompanied by explanatory documentation that remains synchronized as the code evolves.
However, the current literature reveals a significant gap: no papers in the surveyed corpus explicitly describe systems that combine AI agents with literate programming principles. While some systems generate documentation as a separate artifact (Rasheed et al., 2024), none treat documentation and code as a unified, primary source from which both human-readable explanations and executable implementations are derived.
2.3 Multi-Agent Systems Theory
Multi-agent systems (MAS) theory provides the foundational framework for coordinating multiple autonomous entities to achieve complex goals. In software engineering contexts, MAS architectures enable specialization, parallel processing, and distributed problem-solving (Zhu et al., 2022). Key theoretical concepts include agent roles and responsibilities, communication protocols, coordination mechanisms, and emergent behavior from agent interactions.
Role specialization is a central tenet of effective multi-agent systems. By assigning distinct responsibilities to different agents—analogous to roles in human development teams—systems can leverage specialized capabilities and reduce the cognitive complexity faced by any single agent (Nguyen et al., 2024). Communication protocols define how agents exchange information, request assistance, and synchronize their activities. These protocols range from simple message passing to sophisticated negotiation and consensus mechanisms (Grau et al., 2013).
Coordination mechanisms determine how agents’ individual actions combine to achieve system-level goals. Centralized coordination employs a master agent or orchestrator that delegates tasks and integrates results, while decentralized approaches allow agents to self-organize through peer-to-peer interactions (Phan et al., 2024). Hybrid approaches combine elements of both, using centralized planning with decentralized execution (Nguyen et al., 2024).
The application of MAS theory to software development introduces unique challenges. Unlike traditional MAS domains (e.g., robotics, distributed sensing), software development involves complex artifacts (code, tests, documentation), intricate dependencies, and the need for semantic understanding of both requirements and implementations. Recent work demonstrates that MAS architectures can successfully address these challenges, but optimal coordination strategies, communication patterns, and role definitions remain active areas of investigation (Du et al., 2024), (Ribeiro et al., 2025).
3. Multi-Agent Team Coordination and Task Breakdown
3.1 Role Specialization Architectures
Role specialization has emerged as a dominant pattern in multi-agent software development systems, with empirical evidence demonstrating its effectiveness for complex coding tasks. HyperAgent exemplifies this approach through a centralized architecture featuring a Planner agent for advanced reasoning and task delegation, supported by specialized child agents: Navigator (for code exploration), Editor (for code modification), and Executor (for running and testing code) (Phan et al., 2024). This architecture achieved a 33% success rate on the SWE-Bench Verified dataset and 26% on SWE-Bench Lite, representing state-of-the-art performance on industry-standard benchmarks (Phan et al., 2024).
AgileCoder takes a different approach by mapping traditional Agile software development roles to specialized agents: Product Manager, Scrum Master, Developer, Senior Developer, and Tester (Nguyen et al., 2024). This role-based design achieved state-of-the-art performance on HumanEval (70.53% pass@1) and MBPP (80.92% pass@1) benchmarks, with substantial improvements in code executability on the ProjectDev benchmark (Nguyen et al., 2024). The explicit mapping of agent roles to familiar human roles facilitates understanding and provides a natural framework for task allocation.
AutoDev demonstrates autonomous role execution across the development lifecycle, with agents capable of editing files, building projects, running tests, and managing version control operations (Tufano et al., 2024). The system accepts complex natural language objectives and autonomously implements and verifies solutions, achieving 87.8% Pass@1 on test generation tasks (Tufano et al., 2024). This end-to-end autonomy represents a significant advance over systems requiring human intervention at each development stage.
The literature reveals a consistent pattern: role specialization improves both task success rates and code quality compared to monolithic single-agent approaches. However, the optimal granularity of role decomposition remains an open question. Too few roles may overload individual agents with excessive cognitive complexity, while too many roles introduce coordination overhead and potential communication bottlenecks.
3.2 Coordination Mechanisms and Communication Protocols
Effective coordination mechanisms are critical for multi-agent software development systems. HyperAgent employs an asynchronous Message Queue that enables parallel processing and dynamic load balancing, allowing specialized agents to work concurrently on different aspects of a task (Phan et al., 2024). This asynchronous architecture reduces idle time and improves overall system throughput.
AgileCoder introduces a dual-agent conversation design where agents alternate between instructor and assistant roles, maintaining continuity through a Message Stream that preserves conversation history (Nguyen et al., 2024). This design uses unconstrained natural language for agent communication, leveraging the natural language understanding capabilities of underlying LLMs rather than imposing rigid communication protocols (Nguyen et al., 2024).
Cross-Team Collaboration (CTC) extends coordination beyond single teams to enable multiple orchestrated teams to explore alternative decisions and exchange insights (Du et al., 2024). This architecture improves solution quality by considering diverse approaches and selecting the best outcomes from parallel exploration paths (Du et al., 2024). The CTC framework demonstrates that coordination mechanisms can operate at multiple levels—within teams and across teams—to enhance overall system performance.
A critical challenge identified across multiple studies is the computational cost of LLM-based coordination. Each agent interaction that requires LLM inference incurs latency and API costs, potentially creating bottlenecks in highly interactive workflows (Borghoff et al., 2025). This observation motivates research into coordination substrates that minimize LLM involvement in routine coordination tasks while preserving the flexibility of natural language communication for complex interactions.
3.3 Task Decomposition and Dynamic Planning
Task decomposition—breaking complex software development objectives into manageable subtasks—is fundamental to multi-agent effectiveness. HyperAgent’s Planner agent performs hierarchical task decomposition, creating plans that specify sequences of actions for specialized child agents (Phan et al., 2024). The system iteratively refines plans based on execution feedback, demonstrating adaptive planning capabilities (Phan et al., 2024).
AgileCoder structures task decomposition around Agile sprints, with the Product Manager agent creating a backlog of user stories and the Scrum Master organizing these into sprint plans (Nguyen et al., 2024). This approach provides a familiar framework for task breakdown while enabling dynamic adjustment as development progresses (Nguyen et al., 2024). The Dynamic Code Graph Generator (DCGG) component supports context-aware task decomposition by maintaining a graph representation of code dependencies, allowing agents to understand which components must be modified to implement specific features (Nguyen et al., 2024).
AutoDev demonstrates that agents can autonomously decompose natural language objectives into implementation plans without explicit human-provided task breakdowns (Tufano et al., 2024). The system analyzes requirements, identifies necessary code changes, generates implementations, and creates tests to verify correctness—all through autonomous planning and execution (Tufano et al., 2024; Rondon et al., 2025).
Despite these advances, the literature reveals limitations in current task decomposition approaches. Most systems operate on relatively small, well-defined tasks (e.g., single-function implementations, bug fixes) rather than large-scale feature development spanning multiple modules and requiring architectural decisions. The ability to decompose and coordinate work on enterprise-scale features remains an active research challenge (Györy, 2025). Recent work specifically targets project-level decomposition: Zhao et al. (2025) propose ProjectGen, a multi-agent framework that decomposes whole-project generation into architecture design, skeleton generation, and code-filling stages with iterative refinement and memory-based context management. ProjectGen is paired with CodeProjectEval, a benchmark drawn from 18 real-world repositories averaging 12.7 files and 2,388.6 lines of code per task—substantially larger than typical function-level benchmarks (Zhao et al., 2025).
Giannouris and Ananiadou (2026) provide a closely related decomposition pattern in NOMAD, a multi-agent LLM framework that converts natural language requirements into UML class diagrams. NOMAD separates modelling into concept extraction, relationship comprehension, model integration, PlantUML articulation, and optional validation agents, showing that role specialization is not limited to code-writing workflows but also applies to upstream design artifact generation (Giannouris & Ananiadou, 2026). This is particularly relevant to literate-agentic programming because it treats intermediate design models as explicit, inspectable products of the agent workflow rather than hidden reasoning steps.
3.4 Formal Coordination Substrates
An emerging research direction addresses the efficiency and verifiability of agent coordination through formal coordination substrates. The TB-CSPN (Topic-Based Colored Stochastic Petri Net) architecture separates LLM-based topic extraction from deterministic coordination logic implemented as a Petri net substrate (Borghoff et al., 2025). This separation yields significant efficiency gains: fewer LLM calls, faster processing, and higher throughput compared to architectures where LLMs handle both semantic understanding and coordination decisions (Borghoff et al., 2025).
The TB-CSPN approach enables formal verification of coordination properties—such as deadlock freedom, liveness, and fairness—that are difficult or impossible to guarantee in purely LLM-based coordination systems (Borghoff et al., 2025). By encoding coordination logic in a formal substrate with well-defined semantics, the system provides stronger guarantees about agent interactions and system behavior (Borghoff et al., 2025).
This architectural pattern—separating semantic understanding (LLM) from coordination logic (formal substrate)—represents a promising direction for scalable, verifiable multi-agent systems. However, the approach requires careful design of the interface between LLM-based semantic processing and formal coordination mechanisms, and the generalizability of this pattern across diverse software development tasks remains to be empirically validated.
4. Specification-Driven Design with AI Agents
4.1 Natural Language to Executable Specifications
A key capability of modern agentic systems is translating natural language requirements into executable specifications and implementations. AutoDev accepts complex natural language objectives and autonomously assigns agents to implement and verify them, demonstrating a pipeline from specification to execution within an agentic framework (Tufano et al., 2024). This capability reduces the barrier between high-level intent and working software, potentially accelerating early-stage development.
Nagvekar (2025) describes agentic AI-driven CI/CD pipelines where business owners specify goals using natural language, triggering code and test agents to build and deploy applications automatically. This vision of specification-driven automation extends beyond individual coding tasks to encompass entire delivery pipelines (Nagvekar, 2025). However, the paper provides limited empirical validation of this approach in production environments.
Zhao et al. (2025) take a complementary approach, introducing the Semantic Software Architecture Tree (SSAT) as a structured intermediate representation that bridges user requirements and source code. The SSAT is constructed during the architecture-design stage and consumed by downstream skeleton-generation and code-filling agents, providing a semantically rich substrate for cross-agent coordination on project-level tasks (Zhao et al., 2025). This approach addresses the semantic gap between human-written requirements and machine-interpretable structures by making architectural intent explicit and queryable across the multi-agent pipeline.
NOMAD further strengthens the case for explicit design intermediates by focusing on requirements-to-model transformation rather than requirements-to-code transformation. Its pipeline produces UML class diagrams from natural language requirements through a structured JSON representation and PlantUML output, demonstrating a pathway from informal requirements to machine-checkable design artifacts (Giannouris & Ananiadou, 2026). For the dissertation, this supports the argument that a literate-agentic harness should preserve and update design-level representations, not merely generate implementation code and tests.
The challenge of natural language specification lies in ambiguity, incompleteness, and the gap between user intent and technical implementation details. While LLMs demonstrate impressive capabilities in interpreting natural language, they can misinterpret requirements, make unwarranted assumptions, or generate implementations that satisfy the literal specification but miss the underlying intent. Current systems address this through iterative refinement, test-driven validation, and human-in-the-loop verification, but fully autonomous specification interpretation remains an open challenge.
4.2 Planning and Upfront Design Approaches
Planning agents that manage the development lifecycle from conception through verification represent a key component of specification-driven workflows. HyperAgent’s Planner agent performs upfront analysis of tasks, creates implementation plans, and coordinates specialized agents to execute those plans (Phan et al., 2024). This planning phase enables the system to consider multiple approaches, identify potential challenges, and allocate resources effectively before beginning implementation (Phan et al., 2024; see also He et al., 2024, for a broader survey of planning approaches in LLM-based multi-agent SE systems).
AgileCoder embeds Agile planning practices into agent workflows, with sprint planning sessions that prioritize features, estimate effort, and allocate work to developer agents (Nguyen et al., 2024). This structured planning approach provides a process framework for iterative, specification-guided delivery rather than purely reactive code generation (Nguyen et al., 2024).
Wang et al. (2025) identify planning as one of three central capabilities (alongside tool integration and context management) required for effective agentic programming. Their survey highlights diverse planning approaches, from simple sequential task lists to sophisticated hierarchical planning with contingency handling and dynamic replanning (Wang et al., 2025).
Despite these advances, the literature reveals a gap in formal specification enforcement and bidirectional traceability. While systems demonstrate planning capabilities, few provide mechanisms to verify that implementations conform to specifications, trace code elements back to requirements, or automatically detect when code changes violate specified constraints. This gap limits the ability to provide strong guarantees about specification adherence.
4.3 Requirements Engineering with Autonomous Agents
Konda (2025) explores autonomous agents in requirements engineering, testing, and deployment, emphasizing how agents use specifications to guide their decisions and adjust based on changes. This work highlights the potential for agents to participate in requirements elicitation, analysis, and validation—traditionally human-intensive activities (Konda, 2025).
Liang (2022) describes AIDA (Autonomous Intelligent Developer Agent), which uses specification-driven design by interpreting requirements via a semantic knowledge graph. AIDA builds abstract structures from requirements, then generates language-rendered implementations, integrating documentation with requirements, structures, and source code (Liang, 2022). The AIDA v0.1 implementation successfully built a knowledge graph from 20 ontologies (over 22,500 triples) and generated correct source code for an exemplar program in approximately 5 minutes (Liang, 2022).
These examples demonstrate that agents can process formal and semi-formal specifications, but the integration of requirements engineering agents with downstream development agents remains limited. Most systems assume requirements are provided as input rather than collaboratively refined through agent-human interaction. The potential for agents to actively participate in requirements clarification, conflict resolution, and stakeholder negotiation represents an underexplored research direction.
4.4 Traceability and Specification Enforcement
Traceability—maintaining explicit links between requirements, design decisions, implementations, and tests—is essential for specification-driven development but remains underaddressed in current agentic systems. While some architectures maintain internal representations of task decompositions and implementation plans, few provide mechanisms for human developers to query these relationships or verify that implementations satisfy specifications.
The literature reveals that formal specification enforcement and bidirectional traceability lack standardized mechanisms across the corpus. Papers demonstrate prototype pipelines but not widely adopted, verifiable specification contracts that ensure implementations conform to stated requirements. This gap is particularly significant for safety-critical or regulated domains where specification compliance must be demonstrable.
Estimation and scheduling fidelity—predicting build times from specifications—is not systematically demonstrated in the literature. While some architectures report throughput and time improvements, none provide calibrated, generalizable build-time estimators that could enable accurate project planning based on high-level specifications. This represents a critical gap for the dissertation’s goal of achieving more accurate build time estimation through explicit specification and task breakdown.
5. Testing Requirements and Test-Driven Development with AI Agents
5.1 Automated Test Generation
Automated test generation has emerged as one of the most successful applications of AI agents in software development. AutoDev achieved 87.8% Pass@1 on test generation tasks in the HumanEval benchmark, demonstrating that agents can synthesize useful tests as part of the development pipeline (Tufano et al., 2024). This high success rate suggests that test generation may be more tractable for current LLM-based agents than general code generation, possibly because test specifications are often more constrained and verifiable.
The Self-Programming Agent (SPA) integrates pytest, coverage analysis, and static analyzers to drive autonomous code refinement, reporting >80% test coverage and improved code quality through iterative refinement loops (Khemani, 2025). SPA uses test results and coverage metrics to identify gaps in implementations, then autonomously generates additional tests or modifies code to improve coverage (Khemani, 2025). This closed-loop approach demonstrates that agents can use testing as a feedback mechanism for continuous improvement.
AgileCoder includes a dedicated Tester agent responsible for writing test suites, creating testing plans, and identifying bugs (Nguyen et al., 2024). The inclusion of testing as an explicit agent role—rather than an afterthought—reflects the importance of testing in achieving reliable agentic code generation (Nguyen et al., 2024). Empirical results show that writing test suites significantly enhanced performance on the ProjectDev benchmark (Nguyen et al., 2024).
Despite these successes, automated test generation faces challenges in creating comprehensive test suites that cover edge cases, error conditions, and integration scenarios. Most reported metrics focus on unit test generation for individual functions, with less evidence for system-level integration tests or tests that validate complex behavioral requirements.
5.2 Test-Driven Development Methodologies
Test-driven development (TDD)—writing tests before implementations—has been explored in the context of both traditional multi-agent systems and modern agentic software development. Collier et al. (2013) applied TDD to agent-based simulation development, demonstrating that TDD can decompose complex agent behavior into incremental, testable units. This early work established that TDD principles are compatible with agent-based development, though the agents in that study were the subject of development rather than the developers themselves (Collier et al., 2013).
Grau et al. (2013) presented a methodology-agnostic testing process for multi-agent systems based on a V-model, incorporating TDD as a testing strategy. The approach progresses from internal agent component testing (unit tests) through agent testing (integration tests) to system testing (groups of agents interacting) and acceptance testing (Grau et al., 2013). A toolkit extending JUnit with ACL message matchers and mock agents supports automated testing and integration validation (Grau et al., 2013).
In contemporary agentic coding systems, TDD principles manifest as test-first workflows where agents generate tests from specifications before implementing functionality. However, explicit TDD workflows are less common than post-hoc test generation. The potential for agents to fully embrace TDD—using failing tests to drive implementation, refactoring with test coverage as a safety net—remains partially realized in current systems.
5.3 Test Coverage and Quality Metrics
Test coverage metrics provide quantitative measures of testing thoroughness. The Self-Programming Agent (SPA) uses coverage analysis as a core component of its refinement loop, targeting >80% test coverage through autonomous test generation and code modification (Khemani, 2025). The system employs AST (Abstract Syntax Tree) transformations to refine code structure while maintaining or improving coverage (Khemani, 2025).
Beyond simple line or branch coverage, quality metrics for agentic testing include mutation testing (verifying that tests detect intentionally introduced bugs), assertion strength (ensuring tests validate meaningful properties), and test maintainability (generating tests that remain valid as code evolves). However, the literature provides limited evidence on these advanced quality metrics in agentic contexts.
A critical challenge is that high test coverage does not guarantee test quality. Agents can generate tests that execute code without validating correctness, achieving high coverage metrics while providing little actual verification. Ensuring that generated tests include meaningful assertions and cover realistic usage scenarios remains an active research problem.
5.4 Test Environment Automation
Test environment automation—provisioning, configuring, and tearing down test infrastructure—is essential for scalable testing but often requires significant manual effort. Christadoss et al. (2025) describe AI-agent-driven test environment setup and teardown for scalable cloud applications, demonstrating that agents can provision, configure, and dismantle cloud test environments autonomously. This automation reduces setup time and operational overhead, enabling more frequent and comprehensive testing (Christadoss et al., 2025).
AutoDev includes environment orchestration capabilities, managing build processes, dependency installation, and test execution environments as part of its autonomous development workflow (Tufano et al., 2024). This integration of environment management with code generation and testing represents a holistic approach to autonomous development.
However, test environment automation introduces security and isolation challenges. Agents that can provision cloud resources or modify system configurations require careful sandboxing and access controls to prevent unintended consequences. The literature addresses these concerns inconsistently, with some systems implementing robust isolation mechanisms while others provide limited discussion of security considerations.
6. System Integration Validation
6.1 Multi-Dimensional Evaluation Frameworks
Traditional software evaluation focuses on functional correctness—whether code produces expected outputs for given inputs. However, agentic systems require multi-dimensional evaluation that considers tool usage, memory management, environment interactions, and behavioral consistency across executions. Akshathala et al. (2025) present an assessment framework for evaluating agentic AI systems that extends beyond binary task completion to measure tool invocations, memory operations, and environment interactions. This framework reveals runtime uncertainties and behavioral deviations that binary success metrics miss (Akshathala et al., 2025).
NOMAD contributes a domain-specific example of this broader evaluation need by introducing an error taxonomy for LLM-generated UML diagrams, distinguishing class, attribute, and relationship errors such as missing elements, extra elements, misclassified relationships, and semantic misrepresentations (Giannouris & Ananiadou, 2026). The paper also reports that adding a verifier agent improves average structural correctness on the Northwind case study for both GPT-4o and DeepSeek V3, although the authors note that fine-grained verifier behavior remains insufficiently characterized (Giannouris & Ananiadou, 2026). This evidence supports the need for validation metrics that evaluate design artifacts and repair behavior, not just final code execution.
The framework emphasizes that evaluation must account for the nondeterministic nature of LLM-based agents. The same input may produce different execution traces across runs due to sampling variability, context window limitations, or changes in agent state. Robust evaluation requires multiple runs, statistical analysis of success rates, and characterization of failure modes (Akshathala et al., 2025).
Otoum et al. (2026) identify evaluation methodology as a critical concern in their systematic literature review of agentic software engineering. They note that many studies report results on small benchmarks or synthetic tasks, with limited validation on real-world, large-scale software projects. Standardized evaluation protocols that measure throughput, correctness under nondeterminism, integration reliability, and economic metrics remain underdeveloped (Otoum et al., 2026).
6.2 Runtime Validation and Behavioral Assessment
Runtime validation—verifying system behavior during execution rather than only at completion—is particularly important for agentic systems that may exhibit unexpected behaviors or make suboptimal decisions during development. Konda (2025) discusses agentic systems that adjust rollout decisions based on runtime observations, demonstrating adaptive behavior in deployment contexts.
Behavioral assessment frameworks must capture not only whether agents complete tasks successfully but how they approach problems, what strategies they employ, and how they recover from errors. For example, does an agent attempt multiple approaches when initial attempts fail? Does it seek additional information when specifications are ambiguous? Does it validate assumptions before proceeding with implementation? These behavioral characteristics influence reliability and robustness but are rarely quantified in current evaluations.
The literature reveals a gap in standardized behavioral assessment methodologies. While some studies provide qualitative descriptions of agent behaviors, few offer systematic frameworks for characterizing and comparing behavioral patterns across different agent architectures or task domains.
6.3 Continuous Integration with Agentic Systems
Continuous integration (CI) and continuous deployment (CD) pipelines provide natural integration points for agentic systems. Nagvekar (2025) proposes agentic AI-driven CI/CD pipelines for autonomous software delivery, where agents automatically build, test, and deploy code changes in response to natural language goals. The vision includes feedback loops and self-learning processes that enable pipelines to improve over time (Nagvekar, 2025).
Wang et al. (2025) discuss agentic CI/CD proposals that describe end-to-end autonomous delivery with feedback loops, but note that practical deployment details and production safety guarantees remain nascent. The integration of agentic systems into production CI/CD pipelines raises questions about reliability, security, and human oversight that current research has not fully addressed.
A key challenge is balancing automation with safety. While fully autonomous CI/CD promises rapid iteration and deployment, it also risks propagating errors or introducing vulnerabilities without human review. Hybrid approaches that use agents for routine tasks while requiring human approval for critical decisions may offer a pragmatic middle ground, but the optimal division of responsibilities remains an open question.
6.4 Security and Isolation Considerations
Security and isolation are critical for safe integration of agentic systems, particularly when agents have access to development environments, cloud resources, or production systems. AutoDev implements sandboxing and guardrails to isolate agent execution and prevent unintended system modifications (Tufano et al., 2024). These mechanisms include restricted file system access, network isolation, and resource limits (Tufano et al., 2024).
Christadoss et al. (2025) address security in the context of test environment automation, emphasizing the need for proper access controls and audit logging when agents provision cloud infrastructure. Maiti (2026) proposes a zero-trust security architecture for autonomous AI in healthcare, highlighting principles applicable to agentic software development: least privilege access, continuous verification, and explicit trust boundaries.
Despite these contributions, security considerations are addressed inconsistently across the literature. Many papers focus on functional capabilities without discussing security implications, and few provide detailed threat models or security validation results. As agentic systems move toward production deployment, rigorous security analysis and robust isolation mechanisms will become increasingly critical.
7. Structured Development Harness and Workflow
7.1 Agile and Sprint-Based Agent Workflows
Agile methodologies provide a natural framework for structuring agentic development workflows. AgileCoder explicitly maps Agile roles and practices to agent architectures, organizing work into sprints with backlog management, sprint planning, development, testing, and review phases (Nguyen et al., 2024). This structure provides familiar process guardrails while enabling autonomous agent execution within each phase (Nguyen et al., 2024).
The sprint-based approach offers several advantages for agentic systems. Sprints provide natural checkpoints for human oversight, allowing developers to review agent-generated work at regular intervals. Sprint planning sessions enable prioritization and resource allocation, ensuring that agents focus on high-value tasks. Sprint retrospectives could potentially be adapted to analyze agent performance and adjust coordination strategies, though this capability is not demonstrated in current systems.
Ribeiro et al. (2025) explore software development using a multi-agent approach, emphasizing collaborative enhancement of development processes. Their work highlights the potential for agents to participate in Agile ceremonies—stand-ups, planning sessions, retrospectives—though the practical implementation of such participation remains largely conceptual (Ribeiro et al., 2025).
7.2 Development Lifecycle Management
Comprehensive development lifecycle management encompasses requirements analysis, design, implementation, testing, deployment, and maintenance. AutoDev demonstrates end-to-end lifecycle management, with agents handling file editing, building, testing, and version control operations autonomously (Tufano et al., 2024). This holistic approach reduces the need for human intervention at each lifecycle stage, potentially accelerating development velocity (Tufano et al., 2024).
HyperAgent manages the development lifecycle through a central Planner that coordinates specialized agents across analysis, planning, feature localization, code editing, and execution/verification phases (Phan et al., 2024). The system iteratively refines implementations based on test results and execution feedback, demonstrating adaptive lifecycle management (Phan et al., 2024).
However, most systems focus on the implementation and testing phases, with less attention to requirements analysis, architectural design, and long-term maintenance. The ability of agentic systems to participate in high-level design decisions, evaluate architectural tradeoffs, and maintain systems over extended periods remains underexplored. This gap is particularly relevant for the dissertation’s focus on upfront specification and design, where agents must engage with abstract design concepts before generating concrete implementations.
7.3 Context Management and Code Retrieval
Effective context management—providing agents with relevant information about codebases, requirements, and development history—is critical for generating appropriate implementations. AgileCoder’s Dynamic Code Graph Generator (DCGG) maintains a graph representation of code dependencies, enabling context-aware code retrieval that improved executability from 23.38% to 57.50% on the ProjectDev benchmark (Nguyen et al., 2024). This substantial improvement demonstrates the importance of providing agents with structured context about code relationships (Nguyen et al., 2024).
Chatlatanagulchai et al. (2025) conducted an empirical study of context files (Agent READMEs) for agentic coding, investigating how structured context documentation influences agent performance. Their findings suggest that well-designed context files can significantly improve agent understanding of codebases and reduce errors (Chatlatanagulchai et al., 2025).
Context management faces scalability challenges as codebases grow. LLMs have finite context windows, limiting the amount of code and documentation that can be provided as input. Retrieval mechanisms must identify the most relevant context for each task, balancing comprehensiveness with context window constraints. Hybrid approaches combining semantic search, dependency analysis, and relevance ranking show promise but require further development and validation.
7.4 Orchestration Platforms and Agent Operating Systems
Orchestration platforms provide infrastructure for deploying, managing, and coordinating agent teams. Koubâa (2025) proposes Agent Operating Systems (Agent-OS) as a blueprint architecture for real-time, secure, and scalable AI agents. The Agent-OS concept envisions a dedicated runtime environment that handles agent lifecycle management, resource allocation, communication, and security (Koubâa, 2025).
Györy (2025) describes SLA-driven orchestration of long-running multi-agent enterprise workflows, introducing formal scheduling, optimization, and monitoring methods for dynamic coordination. This framework enables adaptive agent selection and real-time SLA enforcement, transforming unpredictable AI systems into dependable enterprise solutions (Györy, 2025).
Erol (2025) analyzes contemporary industrial and academic developments in agent-oriented architecture, identifying common patterns and emerging best practices. The analysis reveals a trend toward standardized orchestration platforms that abstract low-level coordination details and provide high-level APIs for agent team configuration (Erol, 2025).
Practitioner tooling repositories also show that the harness landscape is broader and faster-moving than the peer-reviewed corpus captures. The curated Awesome AI Agents 2026 repository catalogues hundreds of agent tools across categories that map directly to harness design choices, including coding agents, agent frameworks, multi-agent orchestration, protocols and standards, observability and evaluation, safety, and governance (Caramaschi, 2026). Its value for this review is not as empirical evidence of effectiveness, but as a market and tooling map for identifying candidate baselines, protocol trends, and implementation components that may not yet appear in formal studies.
Despite these architectural proposals and practitioner ecosystems, production-grade orchestration platforms for agentic software development remain limited. Most research systems implement custom orchestration logic, hindering reproducibility and limiting adoption. The development of standardized, open-source orchestration platforms could accelerate research and facilitate transition from academic prototypes to production systems.
8. Automated Design Documentation Maintenance
8.1 Current State of Documentation Automation
Documentation automation in current agentic systems primarily focuses on generating API documentation, code comments, and README files as separate artifacts. Rasheed et al. (2024) explored AI-powered code review with LLMs, noting the intention to evaluate LLM-generated documentation updates in future work. This indicates awareness of documentation needs but not yet demonstrated integration (Rasheed et al., 2024).
Some systems generate documentation as a byproduct of code generation, producing comments that explain generated code or creating README files that describe project structure. However, these approaches treat documentation as secondary to code, generated after implementation rather than developed in parallel as a primary artifact.
The challenge of documentation maintenance—keeping documentation synchronized with evolving code—is well-recognized but inadequately addressed. Traditional documentation quickly becomes stale as code changes, leading developers to distrust documentation and rely instead on reading code directly. Agentic systems have the potential to automatically update documentation when code changes, but current implementations rarely demonstrate this capability.
8.2 Literate Programming Integration Gap
The literature review reveals a critical gap: explicit integration of literate programming principles with agentic software development is absent from the surveyed corpus. No papers describe systems where agents generate or maintain literate programs—documents that interweave narrative explanation with executable code as a unified primary artifact.
This gap is significant because literate programming offers a potential solution to the documentation synchronization problem. If documentation and code are maintained as a single source, with code extracted from the documented narrative, then updates to either component naturally maintain synchronization. Agents could generate literate programs where design rationale, implementation decisions, and code are presented as a coherent narrative, facilitating human understanding and long-term maintenance.
The absence of literate programming integration in current agentic systems may reflect several factors: (1) literate programming tools and workflows are not widely adopted in mainstream development, making them unfamiliar to researchers; (2) generating coherent narratives that explain code requires different capabilities than generating code itself; (3) existing benchmarks and evaluation metrics focus on code correctness rather than documentation quality or comprehensibility.
The dissertation research has an opportunity to address this gap by designing and evaluating agentic systems that explicitly incorporate literate programming elements, treating documentation as a first-class artifact maintained in synchronization with code.
8.3 Knowledge Representation and Semantic Graphs
Knowledge representation approaches offer a foundation for maintaining rich, structured documentation. Liang’s (2022) AIDA system uses a semantic knowledge graph to represent requirements, design structures, and code relationships. AIDA integrates documentation with requirements, structures, and source code, building a knowledge graph from ontologies and using it to guide code generation (Liang, 2022).
This approach demonstrates that agents can maintain structured representations of design knowledge, not just generate code. The knowledge graph provides a substrate for traceability, enabling queries about why design decisions were made, how requirements map to implementations, and what dependencies exist between components. However, AIDA focuses on the initial development phase, with limited evidence of how the knowledge graph is maintained as code evolves.
Semantic knowledge graphs could serve as a foundation for literate programming integration, providing a structured representation of design knowledge that agents can query and update. Narrative documentation could be generated from the knowledge graph, ensuring consistency between explanations and the underlying design model. This integration remains a promising but unexplored research direction.
9. Outcomes: Build Time Accuracy, Time-to-Market, and Feature Velocity
9.1 Empirical Performance Metrics
Empirical performance metrics reported in the literature focus primarily on task success rates and code quality measures. HyperAgent achieved 33% success on SWE-Bench Verified and 26% on SWE-Bench Lite, representing state-of-the-art performance on industry benchmarks (Phan et al., 2024). On repository-level code generation, HyperAgent achieved 53.33% Pass@5, and on fault localization (Defects4J), it achieved 59.70% Acc@1, delivering 192 correct fixes (Phan et al., 2024).
AgileCoder reported 70.53% pass@1 on HumanEval and 80.92% on MBPP, with substantial improvements in executability on ProjectDev (Nguyen et al., 2024). AutoDev achieved 87.8% Pass@1 on test generation, demonstrating strong automated test synthesis capabilities (Tufano et al., 2024). The Self-Programming Agent (SPA) reported >80% test coverage and improved code quality through autonomous refinement (Khemani, 2025).
These metrics demonstrate that agentic systems can achieve high success rates on well-defined coding tasks and benchmarks. However, benchmark performance may not fully reflect real-world development challenges, where requirements are ambiguous, codebases are large and complex, and tasks span multiple modules with intricate dependencies. Zhao et al. (2025) explicitly critique this unrealistic-dataset problem, observing that existing function-level benchmarks fail to reflect real-world complexity, and propose CodeProjectEval—constructed from 18 real-world repositories with documentation and executable test cases for automatic evaluation—as a more representative project-level evaluation framework. On this harder benchmark, their ProjectGen framework achieves a 57% improvement in passing test cases over baseline project-level approaches on DevBench, and approximately tenfold more passing test cases on CodeProjectEval, illustrating both the benchmark’s discriminative power and the headroom that remains on realistic project-level tasks (Zhao et al., 2025).
9.2 Efficiency and Throughput Improvements
Efficiency metrics focus on processing time, API call reduction, and throughput. HyperAgent reported average processing times of 106-108 seconds per task and cost-effectiveness of $0.45 per task on SWE-Bench Lite (Phan et al., 2024). These metrics provide a baseline for evaluating the economic viability of agentic development.
The TB-CSPN architecture achieved significant efficiency gains through architectural optimization: fewer LLM calls, faster processing, and higher throughput compared to architectures where LLMs handle both semantic understanding and coordination (Borghoff et al., 2025). This demonstrates that architectural choices significantly impact efficiency, with potential for substantial cost reductions through careful system design (Borghoff et al., 2025).
Liang’s (2022) AIDA system built a knowledge graph from 20 ontologies (over 22,500 triples) and generated source code in approximately 5 minutes on a laptop, demonstrating that knowledge-based approaches can achieve reasonable performance on modest hardware (Liang, 2022).
Despite these efficiency metrics, the literature provides limited evidence on build time estimation accuracy—the ability to predict how long development will take based on specifications. While some systems report actual development times, none demonstrate calibrated estimators that could enable accurate project planning. This gap is directly relevant to the dissertation’s goal of achieving more accurate build time estimation through explicit specification and task breakdown.
9.3 Quality and Reliability Outcomes
Quality metrics beyond functional correctness include code maintainability, test coverage, bug density, and adherence to coding standards. AgileCoder demonstrated that incremental development, code review, and writing test suites significantly enhanced performance, suggesting that process structure improves quality (Nguyen et al., 2024). The Dynamic Code Graph Generator improved executability from 23.38% to 57.50%, indicating that context-aware code generation produces higher-quality implementations (Nguyen et al., 2024).
Cross-Team Collaboration (CTC) reported higher development quality and better exploration of decision spaces compared to baselines, supporting the hypothesis that multiple perspectives improve solution quality (Du et al., 2024). This finding aligns with human software development practices, where code review and collaborative design improve outcomes.
However, long-term quality metrics—such as how well agent-generated code can be maintained and extended over months or years—are not addressed in the literature. Most evaluations focus on immediate task completion, with limited follow-up on whether generated code remains comprehensible and modifiable as requirements evolve. This gap is particularly relevant for the dissertation’s focus on documentation maintenance and long-term system evolution.
9.4 Economic and Business Impact
Economic and business impact metrics—time-to-market, developer productivity, cost savings—are discussed conceptually but rarely quantified empirically. Nagvekar (2025) discusses business impact and innovation potential of agentic CI/CD but does not report specific time-to-market improvements or cost-benefit analyses.
Ulfsnes et al. (2024) conducted an empirical study on transforming software development with generative AI, providing insights into collaboration and workflow changes. However, the study focuses on qualitative observations rather than quantitative economic metrics (Ulfsnes et al., 2024).
The literature reveals insufficient evidence to assert that agentic systems reliably produce accurate, generalized build-time estimates from high-level specifications, or that they demonstrably accelerate time-to-market in production environments. While controlled experiments show task completion improvements, translating these gains to real-world economic benefits requires longitudinal studies in production settings—research that has not yet been conducted at scale.
Feature velocity—the rate at which new features can be developed and deployed—is implicitly addressed through throughput metrics, but explicit measurement of feature velocity in production systems is absent. The dissertation research has an opportunity to design experiments that directly measure build time accuracy, time-to-market, and feature velocity in controlled but realistic development scenarios.
10. Research Gaps and Future Directions
10.1 Identified Research Gaps
The literature review identifies several critical research gaps that the dissertation aims to address:
Literate Programming Integration: No current systems explicitly integrate literate programming principles with agentic software development. The potential for agents to generate and maintain literate programs—where documentation and code form a unified, primary artifact—remains unexplored. This gap represents a significant opportunity for improving documentation quality, synchronization, and long-term maintainability.
Build Time Estimation: While systems report task completion times, none demonstrate calibrated build-time estimators that predict development duration from high-level specifications. Accurate estimation requires understanding task complexity, agent capabilities, and potential failure modes—capabilities not yet demonstrated in the literature.
Documentation Maintenance Automation: Current systems generate documentation as a secondary artifact but do not demonstrate automated maintenance of documentation as code evolves. The challenge of keeping design documentation synchronized with implementations—particularly across large, evolving codebases—remains unsolved.
Specification Traceability: Formal mechanisms for bidirectional traceability between specifications, implementations, and tests are absent. While some systems maintain internal task decompositions, few provide human-accessible traceability that enables verification of specification compliance or impact analysis of proposed changes.
Long-Term Evaluation: Most evaluations focus on immediate task completion, with limited longitudinal studies of how agent-generated code performs over extended periods. Questions about maintainability, evolvability, and comprehensibility of agent-generated systems remain largely unaddressed.
Production Deployment: The transition from research prototypes to production-deployed agentic systems is underexplored. Practical concerns including reliability, security, human oversight, and integration with existing development tools and workflows require further investigation.
10.2 Future Research Directions
Based on identified gaps, several promising research directions emerge:
Literate Agentic Programming: Develop agentic systems that generate literate programs, treating documentation as a primary artifact. Investigate whether literate programming improves human comprehension of agent-generated code, facilitates maintenance, and reduces documentation staleness. Design evaluation methodologies that assess documentation quality alongside code correctness.
Predictive Build Time Models: Create models that estimate development time from specifications, task breakdowns, and historical agent performance data. Investigate factors that influence estimation accuracy, including task complexity, specification clarity, and agent coordination overhead. Validate models through controlled experiments and production deployments.
Automated Documentation Synchronization: Design mechanisms that automatically update documentation when code changes, maintaining consistency between narrative explanations and implementations. Explore knowledge graph representations that support both code generation and documentation generation from a unified design model.
Formal Specification Frameworks: Develop formal specification languages and verification tools tailored to agentic development. Investigate how formal specifications can guide agent behavior, enable automated verification of specification compliance, and support rigorous traceability.
Hybrid Human-Agent Workflows: Explore optimal divisions of responsibility between human developers and agents. Investigate when human oversight is most valuable, how to design effective human-agent collaboration interfaces, and how to transition smoothly between autonomous agent execution and human intervention.
Standardized Evaluation Protocols: Establish standardized benchmarks and evaluation methodologies that assess not only functional correctness but also documentation quality, maintainability, security, and long-term evolvability. Develop metrics for build time accuracy, time-to-market, and feature velocity that can be consistently measured across studies.
10.3 Implications for Dissertation Research
The dissertation research is positioned to address several identified gaps through an integrated approach that combines multi-agent coordination, literate programming principles, and structured development workflows. Specific contributions include:
-
Design and implementation of a literate-agentic development harness that treats documentation and code as unified artifacts, with agents generating and maintaining literate programs throughout the development lifecycle.
-
Empirical evaluation of build time estimation accuracy through controlled experiments that measure how well upfront specification and task breakdown enable accurate prediction of development duration.
-
Investigation of documentation maintenance automation by implementing mechanisms that keep design documentation synchronized with evolving code and measuring documentation staleness over time.
-
Development of specification traceability mechanisms that link requirements to implementations and tests, enabling verification of specification compliance and impact analysis.
-
Longitudinal evaluation of agent-generated systems to assess maintainability, evolvability, and comprehensibility over extended periods, addressing the gap in long-term evaluation.
-
Measurement of time-to-market and feature velocity in realistic development scenarios, providing empirical evidence for economic and business impact claims.
By addressing these gaps, the dissertation will contribute both theoretical insights into literate-agentic programming and practical tools and methodologies that advance the state of the art in AI-assisted software development.
11. Conclusion
This comprehensive literature review has synthesized evidence from 61 studies spanning 2013–2026, examining the state of the art in AI agent systems for software development, multi-agent coordination, specification-driven design, testing automation, system integration validation, structured development workflows, and automated documentation maintenance. The review reveals substantial progress in agentic software engineering, with systems demonstrating impressive capabilities in autonomous code generation, test synthesis, and multi-agent coordination.
Key findings include:
Multi-Agent Coordination: Role specialization architectures (HyperAgent, AgileCoder) achieve state-of-the-art performance on industry benchmarks, with success rates of 26-33% on SWE-Bench and 70-81% on HumanEval/MBPP. Formal coordination substrates (TB-CSPN) offer efficiency gains and verifiability. Cross-team collaboration (CTC) improves solution quality through parallel exploration of alternatives.
Specification-Driven Design: Systems demonstrate natural language to implementation pipelines (AutoDev, AIDA) and requirements-to-model pipelines (NOMAD), with planning agents managing development lifecycles. However, formal specification enforcement, bidirectional traceability, and accurate build-time estimation remain underaddressed.
Testing and TDD: Automated test generation achieves high success rates (87.8% Pass@1 for AutoDev), with systems like SPA achieving >80% test coverage through autonomous refinement. Test environment automation reduces operational overhead. TDD methodologies are compatible with agentic development but not yet fully realized.
Integration Validation: Multi-dimensional evaluation frameworks (Akshathala et al.) extend beyond functional correctness to assess tool usage, memory operations, and behavioral consistency, while NOMAD shows how domain-specific error taxonomies and verifier agents can evaluate design artifacts. However, standardized evaluation protocols and long-term validation remain limited.
Structured Workflows: Agile-based agent workflows (AgileCoder) provide familiar process frameworks. Context management (DCGG) significantly improves code quality. Orchestration platforms and Agent Operating Systems are emerging but not yet production-ready.
Documentation Automation: Current systems generate documentation as secondary artifacts but do not maintain synchronization as code evolves. Literate programming integration—treating documentation and code as unified primary artifacts—is absent from the literature, representing a critical research gap.
Outcomes: Empirical metrics demonstrate task success, efficiency gains, and quality improvements in controlled settings. However, evidence for accurate build-time estimation, demonstrable time-to-market acceleration, and sustained feature velocity improvements in production environments is limited.
The dissertation research is well-positioned to address identified gaps through an integrated approach combining multi-agent coordination, literate programming principles, and structured development harnesses. By explicitly incorporating literate programming elements, implementing build-time estimation mechanisms, automating documentation maintenance, and conducting longitudinal evaluations, the dissertation will advance both theoretical understanding and practical capabilities in AI-assisted software development.
The vision of agent teams working within structured development harnesses, generating and maintaining literate programs that keep design documentation synchronized with code, enabling accurate build-time estimation and accelerating feature velocity, represents a promising direction for future research. This review establishes the foundation for that investigation, identifying proven approaches to build upon and critical gaps to address.
References
Akshathala, S., Srinivasan, A., & Patel, R. (2025). Beyond task completion: An assessment framework for evaluating agentic AI systems. arXiv preprint. https://doi.org/10.48550/arxiv.2512.12791
Borghoff, U. M., Rödig, P., Schlichter, J., & Steinmetz, R. (2025). Beyond prompt chaining: The TB-CSPN architecture for agentic AI. Preprints. https://doi.org/10.20944/preprints202507.1294.v1
Caramaschi, H. G. (2026). Awesome AI Agents 2026 [GitHub repository]. https://github.com/caramaschiHG/awesome-ai-agents-2026
Chatlatanagulchai, W., Pornprasit, C., & Tantithamthavorn, C. (2025). Agent READMEs: An empirical study of context files for agentic coding. arXiv preprint. https://doi.org/10.48550/arxiv.2511.12884
Christadoss, P. R. J., Kumar, S. P., & Rajendran, S. (2025). AI-agent driven test environment setup and teardown for scalable cloud applications. Journal of Knowledge Learning and Science Technology, 4(3), 001. https://doi.org/10.60087/jklst.v4.n3.001
Collier, N., Ozik, J., & Tatara, E. (2013). Test-driven agent-based simulation development. In Proceedings of the 2013 Winter Simulation Conference (pp. 1551-1559). https://doi.org/10.5555/2675983.2676177
Du, X., Liu, M., Wang, K., & Wang, H. (2024). Multi-agent software development through cross-team collaboration. arXiv preprint. https://doi.org/10.48550/arxiv.2406.08979
Erol, O. (2025). Agent-oriented architecture: An analysis on contemporary industrial and academic developments. Preprints. https://doi.org/10.20944/preprints202509.2124.v1
Giannouris, P., & Ananiadou, S. (2026). NOMAD: A multi-agent LLM system for UML class diagram generation from natural language requirements. arXiv preprint. https://doi.org/10.48550/arxiv.2511.22409
Grau, A., Franch, X., & Maiden, N. A. M. (2013). A test driven development of MAS. In Proceedings of the European Workshop on Multi-Agent Systems.
Györy, A. (2025). SLA-driven orchestration of long-running multi-agent enterprise workflows. Zenodo. https://doi.org/10.5281/zenodo.17395641
He, J., Treude, C., & Lo, D. (2024). LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision, and the Road Ahead. ACM Transactions on Software Engineering and Methodology. https://doi.org/10.1145/3712003
Khemani, R. (2025). Self-programming AI: Code-learning agents for autonomous refactoring and architectural evolution. Research Square. https://doi.org/10.21203/rs.3.rs-6688473/v1
Knuth, D. E. (1984). Literate programming. The Computer Journal, 27(2), 97-111.
Konda, S. (2025). Agentic AI for software development: Autonomous agents in requirements engineering, testing, and deployment. [Manuscript in preparation].
Koubâa, A. (2025). Agent operating systems (Agent-OS): A blueprint architecture for real-time, secure, and scalable AI agents. TechRxiv. https://doi.org/10.36227/techrxiv.175736224.43024590/v1
Liang, P. (2022). Autonomous intelligent software development. arXiv preprint. https://doi.org/10.48550/arxiv.2208.06393
Maiti, S. (2026). Caging the agents: A zero trust security architecture for autonomous AI in healthcare. [Manuscript in preparation].
Nagvekar, M. (2025). Agentic AI-driven CI/CD pipelines for autonomous software delivery. In Proceedings of the 2025 International Conference on ICT for Business, Industry and Government (pp. 1-6). https://doi.org/10.1109/ictbig68706.2025.11323919
Nasir, M. J., Kallinteris, A., & Kumar, S. (2025). From code generation to AI collaboration: The role of multi-agent systems in software engineering. ResearchGate. https://doi.org/10.13140/rg.2.2.21102.32320
Nguyen, N. D., Doan, T. N., Phan, H. A., & Nguyen, T. N. (2024). AgileCoder: Dynamic collaborative agents for software development based on agile methodology. arXiv preprint. https://doi.org/10.48550/arxiv.2406.11912
Otoum, Y., Wan, Y., & Nayak, A. (2026). Methods and techniques of agentic software engineering: A systematic literature review. IEEE Access, 14, 12453-12478. https://doi.org/10.1109/access.2026.3652325
Phan, H. A., Nguyen, N. D., Doan, T. N., & Nguyen, T. N. (2024). HyperAgent: Generalist software engineering agents to solve coding tasks at scale. arXiv preprint. https://doi.org/10.48550/arxiv.2409.16299
Rasheed, K., Qadir, M. A., Saeed, M., Zia, T., & Lucini, F. R. (2024). AI-powered code review with LLMs: Early results. arXiv preprint. https://doi.org/10.48550/arxiv.2404.18496
Ribeiro, A. M., Silva, V. T., & Choren, R. (2025). Software development using a multi-agent approach. In Proceedings of the Brazilian Workshop on Social Simulation (pp. 37-48). https://doi.org/10.5753/wesaac.2025.37549
Rondon, P., Wei, R., Cambronero, J., Cito, J., Sun, A., Sanyam, S., Tufano, M., & Chandra, S. (2025). Evaluating agent-based program repair at Google. In Proceedings of the 47th IEEE/ACM International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). https://doi.org/10.1109/ICSE-SEIP66354.2025.00038
Sapkota, H., Kumar, M., & Zhang, L. (2025). Vibe coding vs. agentic coding: Fundamentals and practical implications of agentic AI. arXiv preprint. https://doi.org/10.48550/arxiv.2505.19443
Tufano, M., Deng, S., Sundaresan, N., & Svyatkovskiy, A. (2024). AutoDev: Automated AI-driven development. arXiv preprint. https://doi.org/10.48550/arxiv.2403.08299
Ulfsnes, R., Stray, V., & Moe, N. B. (2024). Transforming software development with generative AI: Empirical insights on collaboration and workflow. [Manuscript in preparation].
Wang, Z., Li, Y., Chen, X., & Liu, M. (2025). AI agentic programming: A survey of techniques, challenges, and opportunities. arXiv preprint. https://doi.org/10.48550/arxiv.2508.11126
Zhang, H. (2024). The rise of AI teammates in software engineering (SE) 3.0: How autonomous coding agents are reshaping software engineering. arXiv preprint. https://doi.org/10.48550/arxiv.2507.15003
Zhao, Q., Zhang, L., Liu, F., Cheng, J., Wu, C., Ai, J., Meng, Q., Zhang, L., Lian, X., Song, S., & Guo, Y. (2025). Towards realistic project-level code generation via multi-agent collaboration and semantic architecture modeling. arXiv preprint. https://doi.org/10.48550/arxiv.2511.03404
Zhu, Y., Wang, Z., Chen, J., & Wang, X. (2022). A survey of multi-agent deep reinforcement learning with communication. [Manuscript in preparation].