Skip to content
AI Engineer Career Launch Programme

Capstone Assessment Rubric

This rubric is used at the Month 6 capstone review. It is published in full so employers, mentors and external reviewers can verify a student's work against the same standard the programme uses, by opening the GitHub repository, visiting the live URL, and checking each criterion below.

Each descriptor is an observable statement, not a subjective grade. Reviewers should be able to point to evidence in the repo, deployment, or design document.

Criterion Weak Adequate Strong Excellent
Does it work Live URL is down or core user flow fails; README has no run instructions; no evidence the RAG pipeline, agent loop or MCP tools execute. Primary user flow completes end-to-end at the public URL; repo runs locally following README steps; at least one retrieval query or agent task returns a sensible response. RAG answers include source citations or chunk references; agent tool calls and MCP server invocations succeed for documented scenarios; unsupported inputs return a clear message instead of a crash or hallucinated answer. All scenarios listed in the design document pass at the live URL; p95 latency for the stated use case is recorded; user feedback or correction loop is wired and visible in the repo.
Architecture Single monolithic script with no separation between API, retrieval, LLM calls and persistence; no architecture diagram in the design document. Distinct modules or services for ingestion, embedding, retrieval, generation and API layer; folder structure matches the diagram in the technical design document. Vector store, agent orchestration (LangGraph, CrewAI or equivalent) and MCP servers are separated with explicit interfaces; configuration and secrets are externalised from application logic. Documented trade-offs for the chosen shape (RAG vs fine-tuning, single vs multi-agent, sync vs async); scaling path and observability hooks at service boundaries are described and implemented.
Code quality API keys or tokens committed to the repo; no linting or formatting; unstructured print statements instead of logging; fewer than four meaningful commits across six months. Secrets loaded from environment variables only; consistent formatting; basic try/except on external calls; commit history shows incremental capstone work, not a single upload. Pydantic models or type hints on API and tool schemas; structured logging with request or trace IDs; unit tests for at least retrieval chunking, tool routing or API endpoints. CI pipeline runs lint and tests on pull request; code review comments addressed in history; test coverage reported for critical paths (ingestion, retrieval, agent loop).
Evaluation No evaluation harness, golden dataset or metrics in the repo; no section on evaluation in the design document. Manual test cases documented with expected vs actual outputs; at least ten labelled queries with pass/fail recorded; basic retrieval precision or answer relevance noted. Automated evaluation harness (script or notebook) runnable from the repo; metrics tracked over time (RAGAS, custom faithfulness/relevance scores or agent task success rate); results committed or linked. Evaluation suite runs in CI or on demand with a fixed seed dataset; before/after comparison after a documented change; failure analysis categorises errors (retrieval miss, tool failure, model refusal).
Reliability Application crashes on empty input, rate-limit response or vector DB timeout; no retries, timeouts or fallback behaviour visible in code. LLM and embedding calls use timeouts and at least one retry; invalid input returns a structured error to the user; graceful message when retrieval returns zero chunks. Rate limiting or queue on the API; circuit-breaker or fallback model path documented; health endpoint reports dependency status (DB, vector store, LLM provider). Defined SLO or error budget for the live deployment; alerting or dashboard shows error rate and latency; runbook in the repo describes steps for common failure modes (provider outage, index corruption).
Deployment Runs only on the student's machine. Deployed and reachable at a public URL. Containerised, environment variables handled correctly, reproducible setup. CI runs on every push, health checks, basic auth or access control in place.
Documentation README is empty or default template; no technical design document linked from the repo; live URL not recorded in README or repo description. README covers local setup, required env vars, architecture overview and live URL; completed technical design document submitted and linked; API endpoints or UI flows described. Deployment guide with platform-specific steps (Docker, Render, Railway or cloud); evaluation results summarised with links to harness output; data ingestion and index rebuild documented. A new developer can onboard from docs alone in under two hours; troubleshooting section covers common errors; changelog or ADR records major design decisions during the six months.
Technical explanation Cannot walk through the repo structure or explain why RAG, agents or MCP were chosen; answers contradict the design document or live behaviour. Explains the end-to-end flow from user query to cited answer or agent action; names the embedding model, vector store and LLM with a reason; answers basic trade-off questions without reading slides. Defends alternatives rejected in the design document with concrete reasons (cost, latency, accuracy); discusses observability and security choices with reference to code; identifies at least three known limitations. Handles panel challenges on failure modes, DPDP/privacy handling and cost at scale; proposes credible next steps tied to measured evaluation gaps; demonstrates debugging a live issue or test failure during review.

Students complete the technical design document template alongside this rubric. Each section of that document maps to one or more criteria above.