Nobody is hired for knowing what RAG is. Build systems that produce a measured number, then package and defend them.
You have finished the tutorials and your portfolio is still a folder of notebooks. This course is nine production-shaped builds and the discipline that makes them count, starting from a rule that governs everything after it: every project ships one measured number against a stated baseline. You will build retrieval that answers questions needing two or three hops across related entities and benchmark it by hop count; a multi-agent research system with typed hand-offs, hard spend ceilings, durable state and a step-level trace you can debug from; and a gateway that survives provider failure with circuit breaking, class-aware failover and per-tenant cost attribution, which you will break on purpose on camera. Then the operations layer that almost no candidate has: a judge calibrated against two hundred of your own labels with weighted kappa and three measured bias numbers, a distillation pipeline whose output is a break-even request volume rather than a model, a red-team harness that attacks your own earlier projects on a schedule and holds regressions at zero, a regression pipeline with shadow mode and gated rollout, a forensics layer that tells you which stage actually failed, and a prompt platform where reverting is a config change rather than a deploy. The last two modules are the part nobody writes down: how to package the work so a stranger can evaluate it in two minutes, and how to talk about it so the number lands before the stack does.
Built by Lakshya Kumar
Paste this into any AI chat. Fill in the bracketed parts with your context — you'll get back a straight answer on whether this belongs on your plate.
I'm taking "Nine AI Projects That Get You Hired" — a portfolio course in the Agentic and Applied AI track. Twelve modules: the standard a project must meet, then nine builds (knowledge-graph RAG, multi-agent research, self-healing gateway, calibrated judge, model distillation, red-team harness, regression CI, failure forensics, prompt A/B platform), then packaging and interview preparation. My context: 1. My current role and years of production experience: [describe] 2. AI work I have shipped so far: [none / prototypes / production] 3. The roles I am targeting: [retrieval / platform / evaluation / safety / agents] 4. Hours per week I can actually commit: [number] 5. What already exists in my portfolio: [describe] Given that, answer: - Which two projects should I build first, and why those two together? - Which project would be a waste of time for the roles I am targeting? - What is a realistic timeline at my stated hours, and where will I underestimate? - For my strongest existing project, what single number is missing that would make it credible?
We grant free access case-by-case — students, career-switchers, builders on a tight budget. Sign in to send us a note.
Sign in to applyFinished the tasks? Take the prompt to your AI and get tested on it. We copy the prompt and open the app — just paste it in.
Retrieval that answers questions needing two or three hops across related entities, benchmarked by hop count against a vector-only baseline.
Supervisor and worker agents with typed hand-offs, hard budgets, durable state in Redis, and a step-level trace you can debug from.
One service in front of every model call: per-provider circuit breaking, class-aware failover, hedging, deferrable queues, and per-tenant cost attribution.
A judge calibrated against 200 human labels with weighted kappa, three measured bias numbers, and CI gates justified by that agreement.
Transfer one narrow task to an 8B student, benchmark quality, cost and latency honestly, and publish the break-even request volume.
A scheduled harness that probes your own earlier projects across eight attack categories, files findings with reproductions, and holds regressions at zero.
Versioned prompts, a hand-built golden set, run diffing with measured significance thresholds, drift detection, shadow mode and a gated rollout.
Stage-level tracing, a failure taxonomy you can count, deterministic replay, failure clustering, and a loop back into the golden set.
A registry with content hashes, deterministic variant assignment, metrics declared up front, automatic guardrail halts and rollback without a deploy.
The README order that gets read, trade-offs with their costs stated, a one-command demo, honest cost figures, and cross-links that turn six repos into one body of work.
Openings built on the problem and the result, one prepared failure per project, resume lines that survive scrutiny, and matching the project to the role.
Read before module 5.