Capstone Assessment Rubric
This rubric is used at the Month 6 capstone review. It is published in full so employers, mentors and external reviewers can verify a student's work against the same standard the programme uses, by opening the GitHub repository, visiting the live URL, and checking each criterion below.
Each descriptor is an observable statement, not a subjective grade. Reviewers should be able to point to evidence in the repo, deployment, or design document.
| Criterion | Weak | Adequate | Strong | Excellent |
|---|---|---|---|---|
| Does it work | Live URL is down or core user flow fails; README has no run instructions; no evidence the RAG pipeline, agent loop or MCP tools execute. | Primary user flow completes end-to-end at the public URL; repo runs locally following README steps; at least one retrieval query or agent task returns a sensible response. | RAG answers include source citations or chunk references; agent tool calls and MCP server invocations succeed for documented scenarios; unsupported inputs return a clear message instead of a crash or hallucinated answer. | All scenarios listed in the design document pass at the live URL; p95 latency for the stated use case is recorded; user feedback or correction loop is wired and visible in the repo. |
| Architecture | Single monolithic script with no separation between API, retrieval, LLM calls and persistence; no architecture diagram in the design document. | Distinct modules or services for ingestion, embedding, retrieval, generation and API layer; folder structure matches the diagram in the technical design document. | Vector store, agent orchestration (LangGraph, CrewAI or equivalent) and MCP servers are separated with explicit interfaces; configuration and secrets are externalised from application logic. | Documented trade-offs for the chosen shape (RAG vs fine-tuning, single vs multi-agent, sync vs async); scaling path and observability hooks at service boundaries are described and implemented. |
| Code quality | API keys or tokens committed to the repo; no linting or formatting; unstructured print statements instead of logging; fewer than four meaningful commits across six months. | Secrets loaded from environment variables only; consistent formatting; basic try/except on external calls; commit history shows incremental capstone work, not a single upload. | Pydantic models or type hints on API and tool schemas; structured logging with request or trace IDs; unit tests for at least retrieval chunking, tool routing or API endpoints. | CI pipeline runs lint and tests on pull request; code review comments addressed in history; test coverage reported for critical paths (ingestion, retrieval, agent loop). |
| Evaluation | No evaluation harness, golden dataset or metrics in the repo; no section on evaluation in the design document. | Manual test cases documented with expected vs actual outputs; at least ten labelled queries with pass/fail recorded; basic retrieval precision or answer relevance noted. | Automated evaluation harness (script or notebook) runnable from the repo; metrics tracked over time (RAGAS, custom faithfulness/relevance scores or agent task success rate); results committed or linked. | Evaluation suite runs in CI or on demand with a fixed seed dataset; before/after comparison after a documented change; failure analysis categorises errors (retrieval miss, tool failure, model refusal). |
| Reliability | Application crashes on empty input, rate-limit response or vector DB timeout; no retries, timeouts or fallback behaviour visible in code. | LLM and embedding calls use timeouts and at least one retry; invalid input returns a structured error to the user; graceful message when retrieval returns zero chunks. | Rate limiting or queue on the API; circuit-breaker or fallback model path documented; health endpoint reports dependency status (DB, vector store, LLM provider). | Defined SLO or error budget for the live deployment; alerting or dashboard shows error rate and latency; runbook in the repo describes steps for common failure modes (provider outage, index corruption). |
| Deployment | Runs only on the student's machine. | Deployed and reachable at a public URL. | Containerised, environment variables handled correctly, reproducible setup. | CI runs on every push, health checks, basic auth or access control in place. |
| Documentation | README is empty or default template; no technical design document linked from the repo; live URL not recorded in README or repo description. | README covers local setup, required env vars, architecture overview and live URL; completed technical design document submitted and linked; API endpoints or UI flows described. | Deployment guide with platform-specific steps (Docker, Render, Railway or cloud); evaluation results summarised with links to harness output; data ingestion and index rebuild documented. | A new developer can onboard from docs alone in under two hours; troubleshooting section covers common errors; changelog or ADR records major design decisions during the six months. |
| Technical explanation | Cannot walk through the repo structure or explain why RAG, agents or MCP were chosen; answers contradict the design document or live behaviour. | Explains the end-to-end flow from user query to cited answer or agent action; names the embedding model, vector store and LLM with a reason; answers basic trade-off questions without reading slides. | Defends alternatives rejected in the design document with concrete reasons (cost, latency, accuracy); discusses observability and security choices with reference to code; identifies at least three known limitations. | Handles panel challenges on failure modes, DPDP/privacy handling and cost at scale; proposes credible next steps tied to measured evaluation gaps; demonstrates debugging a live issue or test failure during review. |
Students complete the technical design document template alongside this rubric. Each section of that document maps to one or more criteria above.