Governed Memory-Driven Multi-Agent Software Development: A Contract-First Methodology for AI-Orchestrated SDLC
Large Language Models (LLMs) have progressed from code-completion assistants to autonomous agents that can plan, invoke tools, edit code, and iterate toward a goal. Yet ad-hoc agentic use rarely changes delivery outcomes at the organizational level: context is lost between sessions, outputs drift, and nothing prevents an agent from acting before the appropriate human approval. We present a governed, memory-driven methodology for agent-orchestrated software development — which we call Claude-Based Development (CBD) — in which a family of specialized LLM agents, each owning a single lifecycle phase, collaborate to carry a change from a business idea to a reviewed, tested, and merged pull request. Rather than treating the model as a single all-purpose chatbot, the methodology decomposes the software development lifecycle (SDLC) into discrete phases — requirements, architecture, planning, implementation, quality, and deployment — and assigns each phase a purpose-built agent with a narrow contract, its own tools, and its own outputs. CBD contributes three mechanisms that distinguish it from prior multi-agent software frameworks: (i) governance gates enforced outside the model, in the tool-execution layer, which make human approval non-bypassable; (ii) a persistent knowledge base that every agent reads before acting and updates afterward, yielding an enrichment cycle in which output quality compounds across iterations; and (iii) contract-first orchestration with delta analysis, which classifies each requirement as reusable, extendable, or net-new so that agents build only what is genuinely new and reuse what already exists. We describe the architecture and end-to-end workflow, present a worked case study, and propose an evaluation framework with operational metrics and research questions for empirically assessing cycle time, reuse, consistency, and governance compliance. We further report preliminary observations from an industrial deployment — a four-tier, four-repository enterprise system with a 75-item backlog — in which delta analysis classified 81% of tier-level work as reusable and in which the methodology's disclosure discipline surfaced a real governance-enforcement gap. We conclude with threats to validity, limitations, and directions for future work.
Authors
- Abhishek Negi
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23022306
- Primary Topic
- Multi-Agent Systems and Negotiation
- Type
- preprint