Embed with a client and land AI into how they run their datacenter, network and workplace — assess, frame, deliver, integrate and govern. 10 modules for infrastructure people who guide and build, not just advise.
This file is generated from the course data by scripts/build-notes.mjs. Edit the course data, not this file.
Organisation: Cranfield Group
Your client for the next ten weeks: Cranfield Group, a mid-market distribution and light- manufacturing firm — two owned datacenters plus a colo they are half-migrated into, an SD-WAN across 42 sites and 5 plants, and roughly 6,000 endpoints. You are embedded from an infrastructure-support practice to make AI actually land in how Cranfield runs the estate — not to sell a platform. Dilan, the Head of Infrastructure, was sold an "AIOps platform" two years ago that became shelfware, and will believe outcomes, not slides. Your own delivery lead wants a repeatable land-and-expand play out of this. Everything you build is one accumulating engagement, defended to the room in Week 10.
Guiding question: What is a forward deployed engineer on an infrastructure estate — and how do I earn a sceptical sponsor's trust?
Outcome: Define the FDE model and separate it from consulting, pre-sales, and traditional MSP break-fix. Explain forward deployment (embed, iterate, earn trust) versus one-shot tool implementation. Draw the co-pilot vs autopilot line for infrastructure specifically, and establish a shared AI-in-operations vocabulary and the FDE duty of care.
Estate lens: You are the person the infrastructure-support firm embeds with a client to make AI actually land in how they run the estate — credible enough to assess it, hands-on enough to build the integration, disciplined enough to be trusted near production. Your deliverable is not a deck; it is a working, governed change that outlasts you. Dilan has been sold "AIOps" before and got shelfware. Trust, earned in weeks, is the whole job.
Apply-at-work mission — Engagement charter + sealed baseline: Write the Cranfield engagement charter: your mission, the trust you must earn from Dilan, and what "success by week 10" looks like. Add a stakeholder power/interest map of the cast. Then seal a baseline: in under 150 words, your current unassisted answer to "what can AI realistically do for infrastructure operations, and what can it not?" — Module 10 reopens it.
Reflection: Where did my mental model of "AI for infrastructure" turn out to be marketing rather than mechanism? Whose trust on this client will be hardest to earn, and why?
Day one on the Cranfield account. The kickoff is done; the goodwill is thin. Dilan opened with "the last lot sold us a dashboard nobody opened" and gave you ten weeks. Your inbox already has a competing appliance quote forwarded by Tom, an eager "can AI just auto-resolve tickets?" from Priya, and a one-line note from Ken: "before anything connects to our data, I want to know what and where."
| Ref | Priority | From | Request |
|---|---|---|---|
| CRN-0001 | P1 | Dilan Petrou | Sponsor kickoff: "prove this won't be shelfware like last time" |
| CRN-0002 | P1 | you, to yourself | Write the engagement charter + who actually decides what |
| CRN-0003 | P1 | Ken Adeyemi | "What data will this touch, and where does it go?" — before any connection |
| CRN-0004 | P2 | Tom Böttcher | Competing appliance quote forwarded — "is this the easy button?" |
Write the Cranfield engagement charter: your mission, the specific trust you must earn from Dilan (who was burned by shelfware), and what "success by week 10" looks like as observable outcomes. Add a power/interest stakeholder map of the six-person cast — who decides, who can block, who to win first. Then seal a baseline: in under 150 words, your current unassisted answer to "what can AI realistically do for infrastructure operations, and what can it not?" Date it and set it aside — Module 10 reopens it.
Deliverable: fde/m01-charter-and-baseline.md — the charter, the stakeholder map, and the dated sealed baseline. This is the first component of your accumulating engagement pack.
Assessment rubric
| Criterion | Weight | What good looks like |
|---|---|---|
| Outcome-framed charter | 30% | Success is stated as observable Week-10 outcomes Dilan would accept, not activities or deliverable counts. |
| Honest stakeholder map | 30% | Each of the six cast members placed on power/interest with a specific reason; names who to win first and why. |
| Genuine sealed baseline | 25% | An unassisted, dated explanation written before the course changes it — its errors are the point, not polish. |
| FDE framing | 15% | Frames the engagement as a governed change that outlasts you, not a demo or a report. |
Drill 1. Kickoff. Dilan opens with "the last lot sold us a dashboard nobody opened" (CRN-0001) and gives you ten weeks.
Task: Write the two-sentence reply that neither dismisses the past failure nor over-promises, and names the one thing you'll do differently. Then state the single Week-10 outcome you'd put your name against.
Drill 2. Ken (CRN-0003): "Before anything connects to our data, I want to know what and where."
Task: List the three questions you must answer for Ken before any tool touches Cranfield data, and say why each matters more than the demo he hasn't asked for.
Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.
1. The clearest difference between a forward deployed engineer and a traditional consultant is that the FDE…
2. Forward deployment differs from a normal SaaS/tool implementation mainly because it…
3. For an infrastructure FDE, the "co-pilot vs autopilot" line is about…
4. Dilan was "burned by shelfware" before. The main lesson for your engagement is that…
5. An LLM sounds fluent and confident even when wrong because…
6. "AI theatre" on an infrastructure account most often looks like…
7. The FDE's "duty of care" on a client estate primarily means…
8. The strongest first move with a sceptical sponsor like Dilan is to…
9. A stakeholder power/interest map is useful because it…
10. Which is genuinely NOT part of the infrastructure FDE role as this course frames it?
11. The sealed baseline you write in Module 1 exists so that…
12. Treating the engagement as "a governed change that outlasts you" mainly implies…
Guiding question: Before I recommend anything, what is actually in this estate, where does it hurt, and where does usable data already live?
Outcome: Run a structured assessment across datacenter, network and workplace. Inventory the data sources that matter — telemetry, event/alert streams, tickets, logs, configs, topology, CMDB — and rate each for AI-readiness. Establish a measured baseline (MTTR, alert volume, ticket mix, capacity headroom) you can later prove improvement against.
Estate lens: Discovery for infrastructure is not a workshop of opinions; it is reading the estate's own exhaust. The CMDB is ~70% true, the monitoring is a three-tool patchwork, and the pain everyone complains about is not always where the data is. Find both. The baseline you take this week is the honest yardstick your Module-10 outcomes get measured against — take it properly or the whole engagement's "impact" is a feeling.
Apply-at-work mission — Estate map, data inventory & baseline: Build the Cranfield estate + data inventory: what telemetry, tickets, configs and CMDB data exist, the trust level of each, and where the CMDB lies. Record a measured baseline for the pains you will target. Mark every unknown with a "?" rather than a guess — the gaps are the engagement. 🎯 Completes Capstone Milestone 1.
Reflection: Which "obvious" pain point turned out to have no usable data behind it? Which data source surprised me by being richer than expected?
Discovery week. The CMDB export arrived and roughly a third of it is wrong. Monitoring is three tools that disagree. Everyone "knows" what hurts most, and no two people agree. You need the estate's own data, not its opinions — and a baseline you can be measured against in Week 10.
| Ref | Priority | From | Request |
|---|---|---|---|
| CRN-0011 | P1 | you, to yourself | CMDB export is ~70% trustworthy — reconcile against reality |
| CRN-0012 | P2 | Marisol Vega | Which of SolarWinds / Datadog / Zabbix is the source of truth? |
| CRN-0013 | P1 | you, to yourself | Baseline the pain: MTTR, alert volume, ticket mix, capacity headroom |
| CRN-0014 | P2 | Ken Adeyemi | "Where does the telemetry and ticket data actually live?" |
Build Cranfield's current-state estate map across datacenter, network and workplace, and a data inventory: for each source (telemetry, event/alert streams, tickets, logs, configs, CMDB) record what it is, how trustworthy it is, and rate it for AI-readiness. Then record a measured baseline for the pains you expect to target — MTTR, alert volume, ticket mix, capacity headroom — the honest yardstick Module 10 measures against. Mark every unknown with a "?" rather than a guess.
Deliverable: fde/m02-estate-map-and-baseline.md — the estate map, the rated data inventory, and the dated baseline with question marks preserved. Second component of the engagement pack.
Assessment rubric
| Criterion | Weight | What good looks like |
|---|---|---|
| Cross-estate coverage | 25% | Datacenter, network and workplace all mapped, with the systems of record (ServiceNow, the monitoring patchwork, CMDB) located. |
| Honest data-readiness rating | 30% | Each source rated for trust and AI-readiness with a reason; the CMDB's unreliability is confronted, not assumed away. |
| Measured baseline | 30% | Concrete before-numbers captured for the targeted pains, dated, so Week-10 improvement can be proven honestly. |
| Question marks preserved | 15% | Unknown cells marked unknown rather than plausibly filled — the gaps are the discovery output. |
Drill 1. The CMDB export lands and ~30% is wrong (CRN-0011): decommissioned hosts present, live services missing.
Task: Describe how you'd reconcile it against something more trustworthy in a week, and name the one class of error you'd check first before trusting any of it for a pilot.
Drill 2. Three monitoring tools disagree on whether DC-1 is healthy (CRN-0012); Marisol asks which to believe.
Task: Give the two-line rule you'd use to decide the source of truth per signal, and one question that would settle it for the specific case of DC-1 capacity.
Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.
1. Discovery on an infrastructure estate should rely most on…
2. Cranfield's CMDB is about 70% trustworthy. The right response is to…
3. The point of taking a measured baseline in Module 2 is to…
4. The "four golden signals" of monitoring are most useful to an FDE for…
5. Before any Cranfield telemetry or ticket data reaches a model, you must…
6. A data source that is high-volume but low-trust (like a noisy alert feed) should be…
7. The most useful thing to do with a data source you cannot yet assess is to…
8. Symptom-versus-cause matters in discovery because the loudest pain…
9. Cranfield runs SolarWinds, Datadog and Zabbix that disagree. For discovery this means…
10. The estate map is most valuable to the engagement as…
11. The biggest risk of skipping a proper baseline is that…
12. When Ken asks "where does the telemetry and ticket data actually live?", the honest FDE answer is…
Guiding question: Across the estate, where does AI genuinely earn its place — and what is each opportunity worth against its risk?
Outcome: Build the infrastructure-AI opportunity catalogue (AIOps and event correlation, observability and anomaly, capacity forecasting, config and change automation, ticket deflection and knowledge, security operations, DC energy). Score each on value × data-readiness × risk, set a quantified KPI hypothesis for the shortlist, and make the buy-vs-integrate-vs-build call. Kill the solution-first ideas early.
Estate lens: The fastest way to become the next shelfware is to demo a clever AI thing nobody needed. The FDE's core judgement is triage: which of the estate's real pains has both high value and enough usable data to be provable, and should you buy a tool, integrate one, or build. A quantified KPI hypothesis — "cut DC-1 alert noise 40% without hiding a real incident" — is what turns a good idea into something you can be held to.
Apply-at-work mission — Opportunity map + KPI hypotheses: Score eight candidate opportunities for Cranfield on value × data-readiness × risk. Shortlist three, and for each state a quantified KPI hypothesis measured against your Module-2 baseline, a risk tier, and a buy/integrate/build recommendation with rationale.
Reflection: Which opportunity was I tempted by because it was interesting rather than valuable? Which "boring" one scored highest, and why?
Framing week. Now the pet ideas surface: Priya wants auto-resolve, a plant manager wants "predictive everything", and the vendor swears their appliance does it all. Tom wants to know which of these is worth money. Your job is triage — value against risk against whether the data even exists.
| Ref | Priority | From | Request |
|---|---|---|---|
| CRN-0021 | P1 | you, to yourself | Rank the AI opportunities across DC / network / workplace |
| CRN-0022 | P3 | Operations | "Predictive everything" request from the plant — real or hype? |
| CRN-0023 | P1 | Tom Böttcher | Tom: "which two of these actually save money, and how much?" |
| CRN-0024 | P2 | Priya Raman | Priya: "can we at least start with password/VPN deflection?" |
Score eight candidate AI opportunities for Cranfield across datacenter, network and workplace on value × data-readiness × risk, using the Module-2 inventory as the readiness input. Shortlist three. For each of the three, state a quantified KPI hypothesis measured against your Module-2 baseline (e.g. "cut DC-1 alert noise 40% without suppressing a real incident"), assign a risk tier, and make a buy / integrate / build recommendation with a one-line rationale. Explicitly name one solution-first idea you killed and why.
Deliverable: fde/m03-opportunity-map.md — the scored catalogue, the shortlist of three with KPI hypotheses and risk tiers, the buy/integrate/build calls, and the killed idea. Third component of the engagement pack.
Assessment rubric
| Criterion | Weight | What good looks like |
|---|---|---|
| Honest scoring | 30% | Each opportunity scored on value, data-readiness and risk with a stated reason; readiness ties back to the Module-2 inventory, not optimism. |
| Quantified KPI hypotheses | 30% | Each shortlisted item has a specific, measurable target tied to the Module-2 baseline — a number you could be held to, not a vibe. |
| Buy/integrate/build judgement | 25% | A defensible recommendation per shortlisted item, with integrate as the reasoned default and build justified only where nothing fits. |
| A killed idea | 15% | Names a genuinely tempting solution-first idea and explains why value or data killed it — evidence of triage, not enthusiasm. |
Drill 1. Tom forwards a competing appliance quote and asks "which two of these actually save money, and how much?" (CRN-0023).
Task: Pick the two Cranfield opportunities you would put in front of Tom, and for each give the one number you would commit to against the Module-2 baseline. Say what you would NOT promise.
Drill 2. Priya asks "can we at least start with password/VPN deflection?" (CRN-0024) while a plant wants "predictive everything".
Task: Rank these two on value × data-readiness × risk in three sentences, and say which you shortlist first and why.
Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.
1. The core triage of value × data-readiness × risk exists mainly to…
2. The fastest route to becoming the next shelfware is to…
3. A quantified KPI hypothesis matters because it…
4. The plant's "predictive everything" request is best treated as…
5. You would lean toward BUYING a tool when…
6. You would consider BUILDING (the rare case) only when…
7. A high-value opportunity with no usable data behind it should be…
8. Ticket deflection scored high for Cranfield mainly because it is…
9. Security-operations AI warrants a heavier risk tier because…
10. Scoring opportunities against the Module-2 baseline matters because…
11. The right output of the triage step is…
12. Killing a solution-first idea early is best read as a sign of…
Guiding question: How do I turn a chosen opportunity into a blueprint I can build safely — and prove I am actually ready to start?
Outcome: Produce a solution blueprint: architecture, scope, dependencies, delivery plan, and acceptance criteria. Reason about blast radius and the advisory-before-action ladder for anything touching production. Then pass the readiness gate — explicit, checkable statements of data/access assumptions, evaluation criteria, and risk controls, without which implementation does not begin.
Estate lens: This module is the bridge from consulting conversation to engineering. A learner can understand a client's problem perfectly and still fail in delivery if they never converted it into a blueprint with data contracts, acceptance criteria, and controls. The readiness gate is a real stage-gate, not a formality: if you cannot write your data-access assumptions, eval criteria and risk controls as checkable statements, you are not ready to build — and saying so is the senior move.
Apply-at-work mission — Blueprint + checkable readiness gate: Produce the Cranfield solution blueprint (architecture, scope, dependencies, delivery plan, acceptance criteria) for your top pilot. Then write the readiness gate as checkable statements: data/access assumptions, evaluation criteria, risk controls, and the human-approval path. A colleague should be able to tick each one true or false. 🎯 Completes Capstone Milestone 2.
Reflection: Which assumption did writing it down as a checkable statement expose as wishful? What would break first if I skipped the gate and started building?
Blueprint week. Dilan has approved a shortlist in principle, but Ken will not let anything start without controls, and Marisol wants to see exactly how a change gets proposed and approved. You cannot enter build until the assumptions, evals and risk controls are written down and checkable.
| Ref | Priority | From | Request |
|---|---|---|---|
| CRN-0031 | P1 | you, to yourself | Solution blueprint + delivery plan for the top pilot |
| CRN-0032 | P1 | Ken Adeyemi | Ken: "show me the data-access, eval and risk controls in writing" |
| CRN-0033 | P1 | Marisol Vega | Marisol: "walk me through how an AI-proposed change gets approved" |
| CRN-0034 | P2 | you, to yourself | Readiness gate: can a colleague tick every assumption true or false? |
Produce the Cranfield solution blueprint for your top shortlisted pilot: architecture, scope, dependencies, delivery plan, and acceptance criteria. Reason explicitly about blast radius and place the pilot on the advisory-before-action ladder. Then write the readiness gate as checkable statements grouped into data/access assumptions, evaluation criteria, risk controls, and the human-approval path — each phrased so a colleague can tick it true or false. If any statement cannot yet be ticked true, say so: that is the gate doing its job.
Deliverable: fde/m04-blueprint-and-readiness-gate.md — the blueprint and the readiness gate as checkable statements, with any un-tickable items flagged. Fourth component of the engagement pack; completes Capstone Milestone 2.
Assessment rubric
| Criterion | Weight | What good looks like |
|---|---|---|
| Buildable blueprint | 30% | Architecture, scope, dependencies and delivery plan are concrete enough to build from; acceptance criteria state what "done and good" means. |
| Blast-radius reasoning | 25% | Places the pilot honestly on the advisory-before-action ladder and reasons about what a wrong action could touch. |
| Checkable readiness gate | 30% | Data/access, eval and risk-control statements are written so a colleague can tick each true or false — not aspirational prose. |
| Honest not-ready flags | 15% | Any assumption that cannot yet be ticked true is flagged rather than hidden — the senior move the gate exists to reward. |
Drill 1. Ken: "show me the data-access, eval and risk controls in writing" before anything starts (CRN-0032).
Task: Draft three checkable readiness-gate statements — one data-access, one eval, one risk control — each phrased so Ken can tick it true or false. Flag any you cannot yet tick true.
Drill 2. Marisol: "walk me through how an AI-proposed change gets approved" (CRN-0033).
Task: Describe the advisory-before-action path for one Cranfield change: what the AI proposes, what it must show, who approves, and where the hard stop is.
Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.
1. This module is the bridge from consulting conversation to…
2. The readiness gate is best understood as…
3. Acceptance criteria are useful mainly because they…
4. Reasoning about blast radius before building means asking…
5. The "advisory-before-action" ladder says that near production you should…
6. A data-access assumption belongs in the readiness gate as…
7. Eval criteria should be defined…
8. Writing an assumption down as a checkable statement most often…
9. The human-approval path in the gate specifies…
10. If you cannot write your risk controls as checkable statements, the right move is to…
11. Naming dependencies in the blueprint matters because…
12. Scope in the blueprint is best kept…
Guiding question: What delivery primitives must be in place before I touch any estate — and which estate will I go deepest on?
Outcome: Stand up the reusable machinery every estate lab depends on. Block A (data & access): data sources and contracts, ITSM/CMDB and API access, identity/RBAC and least-privilege service accounts, ticket/event ingestion, runbook anatomy. Block B (safety & control): tool/function calling, evaluations and acceptance criteria, guardrails, observability, human-in-the-loop gates, fallback paths, and common failure modes. Then nominate your primary estate for deeper work.
Estate lens: This is the module the earlier design was missing. Infrastructure people are strong on estates and weak on agentic patterns and evaluation — so if runbook orchestration, tool calling, guardrails, evals and observability first appear only inside the estate labs, the labs collapse under the load. Learn the delivery primitives once, here, as reusable scaffolding; then the estate modules apply them rather than teach them. This is also where you choose the estate you will go deepest on.
Apply-at-work mission — Delivery foundation + primary-estate choice: Produce the Cranfield data/integration/eval/guardrail plan and stand up the shared scaffold: one permissioned tool/API connection, one grounded assistant, one guarded tool call with an eval and a human-in-the-loop gate. Start the production-readiness checklist you will update every later module. Finally, nominate your PRIMARY estate (Datacenter, Network or Workplace) — it gets your deeper lab and more of the Module-9 integration weight.
Reflection: Which delivery primitive did I underestimate — access/RBAC, evals, guardrails, observability, or fallbacks? Why did I pick the primary estate I chose?
Prepare-the-ground week. Before touching any estate you need the plumbing: read-only API access to ServiceNow and monitoring, least-privilege service accounts Ken will actually sign, an eval harness, guardrails, observability, and a human-in-the-loop gate. Decide which estate you go deepest on before the labs begin.
| Ref | Priority | From | Request |
|---|---|---|---|
| CRN-0041 | P1 | Ken Adeyemi | Request least-privilege service accounts + read-only API scopes |
| CRN-0042 | P1 | you, to yourself | Stand up the eval + guardrail + HITL scaffold once, reuse everywhere |
| CRN-0043 | P2 | Renata Cole | Start the production-readiness checklist (living document) |
| CRN-0044 | P2 | you, to yourself | Nominate your PRIMARY estate: Datacenter, Network or Workplace |
Produce Cranfield's data / integration / eval / guardrail plan, then stand up the shared scaffold the estate labs will reuse: one permissioned tool or API connection (read-only, least-privilege), one grounded assistant over real Cranfield content, and one guarded tool call with an eval and a human-in-the-loop gate. Start the production-readiness checklist you will update every later module. Finally, nominate your PRIMARY estate — Datacenter, Network or Workplace — with a one-paragraph rationale; it gets your deeper lab and more of the Module-9 integration weight.
Deliverable: fde/m05-delivery-foundation.md — the data/access/eval/guardrail plan, the described scaffold (tool connection, grounded assistant, guarded call + eval + HITL gate), the started production-readiness checklist, and your primary-estate choice. Fifth component of the engagement pack.
Assessment rubric
| Criterion | Weight | What good looks like |
|---|---|---|
| Data & access plan | 25% | Data sources, contracts, least-privilege service accounts and read-only API scopes are specified concretely enough for Ken to sign. |
| Safety & control scaffold | 30% | A guarded tool call with an eval and an explicit human-in-the-loop gate is designed — the reusable machinery, not a one-off. |
| Production-readiness checklist started | 20% | A living checklist is begun with real items (evals, rollback, observability, HITL, change record) to be updated each later module. |
| Reasoned primary-estate choice | 25% | Names a primary estate with a rationale tied to value, data-readiness and the learner's own development — not an arbitrary pick. |
Drill 1. Ken will sign least-privilege service accounts and read-only API scopes — but wants them specified first (CRN-0041).
Task: Write the access request for one Cranfield integration: which system, read or write, scoped to what, and why that scope and no wider. Name the one thing you are deliberately NOT asking for.
Drill 2. You must stand up the shared scaffold once and reuse it across all three estates (CRN-0042).
Task: List the four reusable primitives you would build first and, for one guarded tool call, describe the eval and the human-in-the-loop gate you would wrap around it.
Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.
1. Why learn the delivery primitives here in Module 5 rather than inside each estate lab?
2. A least-privilege service account is…
3. A data contract mainly defines…
4. An evaluation ("eval") in this context is…
5. A guardrail is best described as…
6. Observability for an AI workflow means…
7. A human-in-the-loop gate is…
8. A fallback path matters because…
9. Learning runbook anatomy is worth it because…
10. Tool/function calling lets a model…
11. Choosing a primary estate after Module 5 means…
12. The production-readiness checklist should be…
Guiding question: How do I make the datacenter's telemetry tell me things instead of just alarming — without hiding a real incident?
Outcome: Apply the M5 primitives to compute, storage, virtualization and cloud ops: event correlation and alert-noise reduction, anomaly detection on metrics, and capacity forecasting. Integrate with the existing monitoring rather than replacing it, and validate that "fewer alerts" did not suppress something real.
Estate lens: The datacenter is the best first estate because it makes telemetry, compute, storage, patching, observability and operational reliability concrete. It is also where alert fatigue is worst, so noise reduction lands immediately — provided you can prove you did not silence a genuine incident. That validation, not the correlation itself, is the FDE move.
Apply-at-work mission — Datacenter lab: correlation, forecast & validation: Guided lab against sample Cranfield DC telemetry: correlate an alert storm down to root events, build a simple capacity forecast for DC-1's power/cooling constraint, and produce a validation note proving no real incident was suppressed. Deliver the DC runbook artifact. (Deeper build if Datacenter is your primary estate; optional stretch: apply to your own estate.)
Reflection: What did correlation catch that I would have missed — and what, if anything, did it nearly hide? How confident am I the forecast is honest?
Datacenter build. DC-1 threw 900 alerts overnight — a handful real, the rest noise — and it is tight on power and cooling with the colo migration only half done. Dilan remembers the last "AIOps" tool that just moved the noise around. Prove yours reduces it without hiding a real incident.
| Ref | Priority | From | Request |
|---|---|---|---|
| CRN-0051 | P1 | Datadog → NOC | DC-1 alert storm: correlate 900 overnight alerts to root events |
| CRN-0052 | P2 | Dilan Petrou | Capacity: will DC-1 power/cooling hold through the migration? |
| CRN-0053 | P1 | Marisol Vega | Prove no real incident was suppressed by the noise reduction |
Guided lab against sample Cranfield datacenter telemetry. Correlate the DC-1 overnight alert storm down to its few root events; build a simple, uncertainty-aware capacity forecast for DC-1's power/cooling constraint through the colo migration; and — the FDE move — write a validation note proving the noise reduction did not suppress a real incident. Integrate with the existing monitoring rather than replacing it. Deliver the datacenter runbook artifact. (Deeper build if Datacenter is your primary estate; optional stretch: apply the method to your own estate.)
Deliverable: fde/m06-datacenter-runbook.md — the correlation result, the capacity forecast with stated uncertainty, the incident-not-suppressed validation note, and the DC runbook. Updates the production-readiness checklist.
Assessment rubric
| Criterion | Weight | What good looks like |
|---|---|---|
| Correlation to root events | 25% | The alert storm is collapsed to a small set of root events with the reasoning shown, not just a lower count. |
| Honest capacity forecast | 25% | The forecast states its assumptions and plausible error range, and addresses DC-1's real power/cooling constraint. |
| Incident-not-suppressed validation | 35% | A concrete check demonstrates no genuine incident was hidden by the noise reduction — the safety proof, not an afterthought. |
| Integration over replacement | 15% | The pilot enriches the existing monitoring patchwork rather than proposing to rip and replace it. |
Drill 1. DC-1 threw ~900 alerts overnight; a handful are real (CRN-0051), and Dilan remembers the last tool that just moved the noise around.
Task: Describe how you would collapse the storm to root events AND the one check you would run to prove a genuine incident was not buried. State what you would show Dilan.
Drill 2. DC-1 is tight on power and cooling with the colo migration half done (CRN-0052).
Task: Sketch a capacity forecast that Dilan could act on, naming its key assumption and the range of error you would attach to it.
Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.
1. Why is the datacenter a good first estate to build in?
2. The real FDE move in alert-noise reduction is…
3. Event correlation aims to…
4. Anomaly detection on metrics is most valuable when…
5. A capacity forecast for DC-1 must above all be…
6. You integrate with existing monitoring rather than replacing it because…
7. The four golden signals mainly help you…
8. The incident-not-suppressed validation note matters because…
9. Correlating the alert storm down to root events means…
10. A capacity forecast is honest when…
11. A real incident hidden by noise reduction is…
12. The datacenter runbook artifact should show…
Guiding question: Can AI find the "random" branch slowdown before the NOC does — and never make an unsafe change?
Outcome: Apply AI to network telemetry for assurance and anomaly across sites, add config-drift detection and pre-change validation, and enable natural-language query over network state — with hard guardrails keeping every change human-approved.
Estate lens: Network changes are the highest-blast-radius action in the estate, and Marisol's change control is usually right to say no. The FDE earns her trust by respecting the CAB, not bypassing it: AI proposes and explains a change; a human disposes. Get that boundary wrong once and you lose the room for the rest of the engagement.
Apply-at-work mission — Network lab: root-cause & human-gated change: Guided lab against sample Cranfield network telemetry and configs: trace the branch-slowdown to a cause using telemetry + natural-language query, detect a config drift, and propose (never auto-apply) a fix with a validation checklist and an explicit human-approval gate. Deliver the network runbook artifact. (Deeper if Network is your primary estate.)
Reflection: Where was I tempted to let the agent act rather than propose? What would have to be true before any network change could be safely automated here?
Network build. The "random" branch slowdown is back — three sites, no root cause after weeks — and a shipping-peak change freeze starts Friday. Marisol will consider an AI assist, but nothing it suggests touches the network without her sign-off.
| Ref | Priority | From | Request |
|---|---|---|---|
| CRN-0061 | P1 | Marisol Vega | Root-cause the "random" slowdown across 3 branches |
| CRN-0062 | P1 | Network ops | Config drift detected pre-freeze — validate before any change |
| CRN-0063 | P1 | Marisol Vega | Every AI-proposed change stays human-approved — show the gate |
Guided lab against sample Cranfield network telemetry and configs. Trace the "random" branch slowdown to a cause using telemetry plus natural-language query; detect a config drift ahead of the shipping-peak freeze; and propose — never auto-apply — a fix with a validation checklist and an explicit human-approval gate. Every AI-proposed change stays human-approved. Deliver the network runbook artifact. (Deeper build if Network is your primary estate.)
Deliverable: fde/m07-network-runbook.md — the root-cause trace, the drift finding, and the proposed fix with its validation checklist and human-approval gate. Updates the production-readiness checklist.
Assessment rubric
| Criterion | Weight | What good looks like |
|---|---|---|
| Root-cause trace | 30% | The branch slowdown is traced to a plausible cause with the telemetry/query reasoning shown, not asserted. |
| Config-drift finding | 20% | A drift from intended config is detected and framed as an early warning before the freeze. |
| Human-gated change design | 35% | The fix is proposed with evidence and an explicit approval gate; nothing is auto-applied — the boundary Marisol trusts. |
| Validation checklist | 15% | A pre-change validation checklist makes the proposed change safe to review and reversible. |
Drill 1. The "random" branch slowdown is back across three sites with no root cause after weeks, and a change freeze starts Friday (CRN-0061).
Task: Outline how you would use telemetry + natural-language query to trace it, and state clearly where the AI stops and a human takes over.
Drill 2. Config drift is detected pre-freeze (CRN-0062) and must be validated before any change.
Task: Write the human-approval gate for the proposed fix: what the AI must show, who approves, and the hard stop if validation fails.
Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.
1. Network changes are best understood as…
2. You earn Marisol's trust by…
3. Config-drift detection is…
4. Pre-change validation means…
5. Natural-language query over network state lets an engineer…
6. AI must never auto-apply a network change because…
7. AI helps root-cause the "random" branch slowdown mainly because…
8. A change freeze before the shipping peak means…
9. The human-approval gate on a network change is…
10. When AI proposes a network fix, it must also…
11. Config drift detected just before a freeze is best treated as…
12. The network runbook artifact should include…
Guiding question: How do I cut the service desk's load and lift end-user experience without leaking data or breaking devices?
Outcome: Apply AI to identity, endpoint and collaboration workflows: digital-employee-experience and endpoint analytics, ticket deflection with retrieval grounded in the client's runbooks, and a bounded self-heal automation — with anonymisation and PII discipline throughout, and honest measurement of what was actually deflected.
Estate lens: Workplace comes last because endpoint and collaboration scenarios depend on identity, access, connectivity and device posture — the things the datacenter and network modules made concrete. It is also the most data-sensitive estate: tickets are full of names and machines, so the sanitisation discipline is not optional. "Deflected" only counts if it was not merely deflected-then-reopened.
Apply-at-work mission — Workplace lab: deflection & self-heal: Guided lab against sample Cranfield runbooks and tickets: build a retrieval-grounded assistant over the runbook set, design one bounded self-heal flow with explicit stop conditions, add the sanitisation step, and define an honest deflection measurement. Deliver the workplace runbook artifact. (Deeper if Workplace is your primary estate.) 🎯 Completes Capstone Milestone 3.
Reflection: What did I have to strip before anything reached a model, and did I nearly miss a quasi-identifier? Is my deflection metric honest, or flattering?
Workplace build. Priya's queue is drowning — 18 password resets, 12 VPN, Outlook, onboarding — and she is ready to pilot deflection today. Ken's only condition: no user data leaks into a model. Measure what is actually deflected, not what merely bounced back later.
| Ref | Priority | From | Request |
|---|---|---|---|
| CRN-0071 | P1 | Priya Raman | Ticket deflection pilot over Priya's runbooks/KB |
| CRN-0072 | P2 | you, to yourself | One bounded self-heal flow with explicit stop conditions |
| CRN-0073 | P1 | Ken Adeyemi | Ken: "prove no PII reaches the model" before go-live |
Guided lab against sample Cranfield runbooks and tickets. Build a retrieval-grounded assistant over Priya's runbook set; design one bounded self-heal flow with explicit stop conditions; add the sanitisation step that strips the identity payload (including quasi-identifiers) before anything reaches a model; and define an honest deflection measurement that counts resolution, not deflected-then-reopened. Deliver the workplace runbook artifact. (Deeper build if Workplace is your primary estate.) Completes Capstone Milestone 3.
Deliverable: fde/m08-workplace-runbook.md — the grounded assistant design, the bounded self-heal flow with stop conditions, the sanitisation step, and the honest deflection metric. Updates the production-readiness checklist; completes MS3.
Assessment rubric
| Criterion | Weight | What good looks like |
|---|---|---|
| Grounded assistant | 25% | Answers come from Cranfield's own runbooks via retrieval, with a clear reason grounding reduces made-up answers. |
| Sanitisation discipline | 30% | The identity payload — including quasi-identifiers, not just names — is stripped before any data reaches the model. |
| Bounded self-heal | 25% | One self-heal flow has explicit stop conditions and a hand-off to a human — bounded, not free-roaming. |
| Honest deflection metric | 20% | Deflection is measured as genuine resolution, explicitly tracking reopened tickets rather than first-touch bounces. |
Drill 1. Priya's queue is drowning and she is ready to pilot deflection today; Ken's only condition is that no user data leaks into a model (CRN-0071, CRN-0073).
Task: Describe the sanitisation step you would put before the model and name one quasi-identifier you would strip that is not an obvious name.
Drill 2. You are designing one bounded self-heal flow with explicit stop conditions (CRN-0072).
Task: Define the flow's single job, its two hardest stop conditions, and how you would measure deflection honestly rather than flatteringly.
Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.
1. The workplace is best understood as the estate that is most…
2. Sanitisation before data reaches a model is…
3. A retrieval-grounded assistant answers from…
4. A ticket only counts as "deflected" if it…
5. A bounded self-heal flow needs above all…
6. A quasi-identifier is dangerous because…
7. The workplace estate is built last because…
8. Grounding an assistant in real runbooks mainly…
9. An honest deflection metric must also track…
10. The sanitisation step should strip…
11. A self-heal flow that hits its stop condition should…
12. The workplace runbook artifact should show…
Guiding question: How do I combine the estate pilots into one production-grade workflow that is monitored, reversible, and safe to run?
Outcome: Integrate the pilots into cross-estate agentic runbooks (reason → propose → human approves → act → verify), connected to ITSM/monitoring/config via APIs and webhooks with least-privilege accounts. Prove production readiness: eval evidence, observability dashboards, rollback, change records, and a handover pack — the M5 foundations, now demonstrated at integration.
Estate lens: This is synthesis, not first exposure — the runbook, tool-calling, eval, guardrail and observability foundations were taught in Module 5 precisely so this module can integrate rather than introduce. The bar here is production-readiness: an agent that touches production must propose-and-verify behind a human gate, log every action, and roll back cleanly. If it cannot, it is a demo, not a delivery.
Apply-at-work mission — Integrated runbook pack + production-readiness review: Wire one pilot to ServiceNow (ticket in → AI triage → drafted action → human approve → verify) and build one cross-estate agentic runbook with an explicit stop-and-approve. Run the production-readiness checklist against it — test evidence, rollback, observability, HITL, change record, handover — and deliver the integrated runbook pack with observability evidence and the HITL gate design. 🎯 Completes Capstone Milestone 4.
Reflection: Which production-readiness item was my runbook missing when I first checked? Where is the single point where a wrong AI step could still cause harm?
Integration week. The three pilots work in isolation; now they must become one production-grade workflow wired to ServiceNow, monitored, and reversible. Ken will not approve anything that acts on production without a human gate, an audit trail, and a clean rollback. This is where "demo" becomes "delivery".
| Ref | Priority | From | Request |
|---|---|---|---|
| CRN-0081 | P1 | you, to yourself | Wire a pilot to ServiceNow: ticket → AI triage → human approve → verify |
| CRN-0082 | P1 | Renata Cole | Run the production-readiness checklist against the runbook |
| CRN-0083 | P1 | Ken Adeyemi | Ken: "show me rollback, audit trail and the human gate" |
| CRN-0084 | P2 | Marisol Vega | Observability: dashboards + alerts for the AI workflow itself |
Wire one pilot to ServiceNow (ticket in → AI triage → drafted action → human approve → verify) with least-privilege access, and build one cross-estate agentic runbook with an explicit stop-and-approve step. Run the production-readiness checklist against it — eval evidence, rollback, observability, HITL gate, change record, handover — and deliver the integrated runbook pack with observability evidence and the HITL gate design. This is synthesis of the Module-5 foundations at integration, not first exposure. Completes Capstone Milestone 4.
Deliverable: fde/m09-integrated-runbook.md — the wired ServiceNow loop, the cross-estate runbook with its stop-and-approve, the completed production-readiness review, and the observability + HITL evidence. Completes MS4.
Assessment rubric
| Criterion | Weight | What good looks like |
|---|---|---|
| Integrated, gated workflow | 30% | One pilot is wired end to end (ticket → triage → approve → verify) with an explicit human gate before any action. |
| Production-readiness evidence | 30% | Eval evidence, rollback, observability and change record are demonstrated against the living checklist, not merely claimed. |
| Least-privilege integration | 20% | API/webhook access is scoped to the workflow with least privilege — the integration Ken would sign. |
| Synthesis, not re-teaching | 20% | The runbook applies the Module-5 primitives at integration and names the single point where a wrong step could still cause harm. |
Drill 1. You must wire a pilot to ServiceNow: ticket in → AI triage → human approve → verify (CRN-0081), and Ken wants rollback, audit trail and the human gate shown (CRN-0083).
Task: Describe the least-privilege access you would request and where exactly the human gate sits. Name the one action in the loop that must never happen without approval.
Drill 2. You are running the production-readiness checklist against the runbook (CRN-0082).
Task: List the readiness items you would check and identify the one your runbook is most likely to be missing on the first pass.
Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.
1. Module 9 is best understood as…
2. An agentic runbook near production follows the pattern…
3. "Production readiness" here means…
4. The integration to ServiceNow should use…
5. The real difference between a demo and a delivery is that…
6. A clean rollback matters because…
7. Observability dashboards for the AI workflow itself let the ops team…
8. A change record for an AI action exists so that…
9. The handover pack matters because…
10. The single most dangerous point in an integrated runbook is…
11. At integration, the Module-5 foundations are…
12. A cross-estate agentic runbook needs an explicit stop-and-approve because…
Guiding question: How do I make this outlast me — governed, adopted by the client's team, measured, and defended to a sceptical room?
Outcome: Consolidate the governance thread into a signable AI-in-operations pack (data residency, change control for AI actions, guardrails, audit, the autonomy ladder). Drive adoption and handover to the client's ops team, measure outcomes and ROI against the Module-2 baseline, tell the demo narrative, plan the land-and-expand, and defend the whole engagement to the Cranfield cast.
Estate lens: An engagement that dies when the FDE leaves was theatre. Landing means the client's own team runs it, the CISO and change board have signed the governance, the CFO has a defensible number, and there is a plan to expand from pilot to estate. The capstone is not a presentation — it is the defence of accumulated work in front of people paid to be sceptical.
Apply-at-work mission — Capstone: defend the Cranfield engagement: Assemble the accumulated engagement pack — charter, estate map + baseline, use-case canvas, blueprint + readiness gate, delivery foundation, the estate runbooks, the integrated runbook pack, and the governance/adoption pack — and defend it to the cast with measured before/after. Choose a track: (A) delivery-lead package (the full engagement, presentation-ready) or (B) a deep single-estate implementation with a working integration.
Reflection: Reopen the Module-1 sealed baseline: what did I get wrong on day one? What would I need to be true before I let any of this run unattended?
Final week. Steering review Friday: Dilan wants to know it will outlast you, Tom wants the number against the Week-2 baseline, Ken wants a signable governance pack, and Renata wants a handover the client's own team can run. Defend the whole engagement — and make the case to expand from pilot to estate.
| Ref | Priority | From | Request |
|---|---|---|---|
| CRN-0091 | P1 | Dilan Petrou | Capstone defense: architecture, evidence, governance, adoption plan |
| CRN-0092 | P1 | Tom Böttcher | Tom: "the number, measured against where we started" |
| CRN-0093 | P1 | Ken Adeyemi | Signable AI-in-operations governance pack |
| CRN-0094 | P2 | Renata Cole | Handover pack + 90-day land-and-expand plan |
Assemble the accumulated engagement pack — charter, estate map + baseline, opportunity map, blueprint + readiness gate, delivery foundation, the three estate runbooks, the integrated runbook pack, and the governance/adoption pack — and defend it to the Cranfield cast with a measured before/after against the Module-2 baseline. Consolidate governance into a signable AI-in-operations pack (data residency, change control for AI actions, guardrails, audit, the autonomy ladder), plan the handover and the 90-day land-and-expand, and reopen the Module-1 sealed baseline. Choose a track: (A) delivery-lead package — the full engagement, presentation-ready; or (B) a deep single-estate implementation with a working integration.
Deliverable: fde/m10-capstone.md — the assembled engagement pack, the signable governance pack, the measured ROI against the Module-2 baseline, the adoption + land-and-expand plan, and the reopened sealed baseline. The defended whole.
Assessment rubric
| Criterion | Weight | What good looks like |
|---|---|---|
| Measured before/after | 25% | Outcomes are shown against the dated Module-2 baseline as a defensible number, not a feeling or a projection. |
| Signable governance pack | 25% | Data residency, AI change control, guardrails, audit and the autonomy ladder are consolidated into something the CISO and CAB could actually sign. |
| Adoption & land-and-expand | 25% | A handover plan that lets the client's team run it, plus a 90-day expansion path sequenced by value and readiness. |
| Honest reflection | 25% | The Module-1 sealed baseline is reopened and what changed is named honestly — the growth, not a polished restatement. |
Drill 1. Steering review Friday: Dilan wants proof it will outlast you, Tom wants the number against the baseline, Ken wants a signable pack (CRN-0091, CRN-0092, CRN-0093).
Task: For each of the three, state the one artifact you would put in front of them and the single sentence you would open with.
Drill 2. Renata wants a handover pack and 90-day land-and-expand plan the client's own team can run (CRN-0094).
Task: Outline the first three expansion steps after the pilot and the one condition that must be true before each, so expansion is evidence-led not hope-led.
Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.
1. An engagement that dies the moment the FDE leaves was…
2. "Landing" the engagement means…
3. The governance pack must be…
4. ROI must be measured…
5. The autonomy ladder is best understood as something…
6. Adoption by the client's team requires…
7. The capstone is best understood as…
8. Change control for AI actions means…
9. Reopening the Module-1 sealed baseline shows…
10. A defensible ROI number is one that…
11. The land-and-expand plan is valuable because…
12. Data residency belongs in the governance pack because…
Unlocks in module 1.
A quick self-assessment across the six competencies, mapped to targeted "do these before Module 5" actions — not just a score.
# Pre-work diagnostic → remediation map
Rate yourself 1–5 on each axis, then follow the remediation for anything under 3.
The point is targeted preparation before the delivery foundation (Module 5), not a grade.
| Competency | 1–5 | If under 3, do this before M5 |
|---|---|---|
| Infrastructure depth (DC / network / workplace) | | Sketch your own estate's topology and systems of record from memory, then check it |
| AI fluency in operations | | Watch the Karpathy intro (M1) and write one paragraph on what an LLM cannot do |
| Discovery & scoping | | Practise turning one real pain into a value × data-readiness × risk line |
| Hands-on delivery | | Read the Anthropic tool-use overview (M5) and the evals piece; note what you don't follow |
| Governance, risk & security | | Read the OWASP LLM Top 10; list the three risks most relevant to your estate |
| Client trust & communication | | Write the two-sentence reply to a sceptical sponsor (M1 drill) |
## How to use it
- Retake it after Module 5 and again at Module 10 — the delta is evidence for your journal.
- Bring the lowest axis into your primary-estate choice: sometimes the right estate is the
one that stretches your weakest competency, sometimes the one that plays to your strongest.
Unlocks in module 3.
The value × data-readiness × risk grid, with a quantified KPI hypothesis and a buy / integrate / build call per shortlisted item.
# Opportunity-scoring sheet
Score each candidate 1–5 on all three axes. Shortlist the highest combined, then commit a
number. Readiness comes from the Module-2 data inventory, not optimism.
| Opportunity | Value (1–5) | Data-readiness (1–5) | Risk (1–5, lower=safer) | Shortlist? |
|---|---|---|---|---|
| DC alert-noise reduction | | | | |
| Capacity forecasting | | | | |
| Network anomaly / assurance | | | | |
| Config-drift detection | | | | |
| Ticket deflection | | | | |
| Bounded self-heal | | | | |
| Security operations assist | | | | |
| DC energy optimisation | | | | |
## For each shortlisted item
- **KPI hypothesis (quantified, vs the M2 baseline):** e.g. "cut DC-1 alert noise 40% without
suppressing a real incident."
- **Risk tier:** low / medium / high — and the one control that earns the tier.
- **Buy / integrate / build:** with a one-line rationale. Integrate is the default; build is
the exception, reserved for specific, valuable needs nothing else fits.
## The killed idea
Name one tempting solution-first idea you rejected, and whether value or data killed it.
Recording the rejection is the evidence of triage.
Unlocks in module 4.
Architecture, scope, dependencies, delivery plan and acceptance criteria — the bridge from a chosen opportunity to a buildable plan.
# Solution-blueprint template
## 1. Problem & outcome
- The bounded problem this pilot solves (one sentence).
- The acceptance criteria: what "done and good" looks like, as checkable statements.
## 2. Architecture
- Data sources in → processing → action out. Where the AI sits and what it may touch.
- The advisory-before-action rung this pilot occupies (advise / propose / act-with-approval).
## 3. Scope
- In scope (this pilot). Out of scope (named later phases).
## 4. Dependencies
- Access grants, data feeds, sign-offs, environments. The unnamed dependency is the one
that stalls the build in week six — write them all down.
## 5. Delivery plan
- Milestones, sequence, who does what, and the human-approval path for any action.
## 6. Blast radius
- What a wrong action could touch, and how far it could spread. The control follows the reach.
Unlocks in module 5.
The reusable data / access / eval / guardrail scaffold every estate lab depends on. Stood up once in M5, referenced everywhere after.
# Delivery-foundation checklist (living — start in M5)
## Block A — data & access
- [ ] Data sources identified with a contract (what, shape, who may use it)
- [ ] ITSM / CMDB / monitoring API access scoped and requested
- [ ] Least-privilege service accounts specified (read-only where possible)
- [ ] Identity / RBAC model written down and signed by the CISO
- [ ] Ticket / event ingestion path defined
- [ ] Runbook anatomy understood for the target workflow
## Block B — safety & control
- [ ] Tool / function calling interface defined and bounded
- [ ] Evals written as acceptance criteria (before build)
- [ ] Guardrails specified (what the AI may and may not do)
- [ ] Observability designed (traces, metrics, logs for the AI itself)
- [ ] Human-in-the-loop gate placed before any production action
- [ ] Fallback path defined (what happens when the AI is unsure or fails)
- [ ] Failure modes reviewed against the failure-modes library
## Primary-estate choice
- [ ] Primary estate nominated (Datacenter / Network / Workplace) with a rationale
Unlocks in module 9.
The go/no-go for letting a runbook near production. Begun in M5, updated every module, run in full at integration (M9).
# Production-readiness checklist (living — run in full at M9)
A runbook that cannot tick these is a demo, not a delivery.
## Evidence
- [ ] Eval evidence: the AI meets its acceptance criteria, shown not claimed
- [ ] Observability: dashboards + alerts for the AI workflow itself
- [ ] Change record: every AI-driven action is auditable after the fact
## Safety
- [ ] Human-in-the-loop gate before any production action
- [ ] Clean rollback for every action the AI can take
- [ ] Least-privilege access, scoped to the workflow
- [ ] The single point where a wrong step could still cause harm is identified and gated
## Handover
- [ ] The client's own team can operate it (runbook + training)
- [ ] Handover pack complete
- [ ] Escalation and stop conditions documented
Unlocks in module 5.
The catalogue of ways AI-in-operations goes wrong — from AI theatre to suppressed incidents to prompt injection — each with its guardrail.
# Failure-modes & anti-patterns library
Referenced in M5 (foundation), M9 (integration), and the governance pack.
## Delivery anti-patterns
- **AI theatre** — a slick demo of an unowned, unneeded capability. *Guard:* every pilot has a
named owner and a measured problem behind it.
- **Solution-first** — building the clever thing, then hunting for a problem. *Guard:* triage on
value × data-readiness × risk before any build.
- **Baseline-free impact** — "it feels faster." *Guard:* a dated Module-2 baseline behind every KPI.
- **The hero engagement** — only the FDE can run it. *Guard:* handover pack + the client's team operates it.
## Operational failure modes
- **Suppressed incident** — noise reduction hides a real event. *Guard:* the not-suppressed validation note.
- **Confident wrong answer** — fluent hallucination. *Guard:* retrieval grounding + evals + a human gate.
- **Unattended production action** — AI acts without approval. *Guard:* HITL gate + advisory-before-action ladder.
- **High-blast-radius change** — one network change takes down many sites. *Guard:* propose-not-apply + CAB.
## Data & security failure modes
- **Identity leak** — names / hosts / IPs reach a model. *Guard:* the sanitisation step (incl. quasi-identifiers).
- **Prompt injection / excessive agency** (OWASP LLM Top 10). *Guard:* input handling + bounded tools + least privilege.
- **Data residency breach** — telemetry leaves its permitted region. *Guard:* mapped data locations + governance pack.
Unlocks in module 10.
The signable AI-in-operations pack the CISO and change board approve — data residency, AI change control, guardrails, audit, the autonomy ladder.
# AI-in-operations governance pack (skeleton)
Threaded from Module 1, consolidated here into something the CISO and CAB can sign.
## 1. Scope & data residency
- Which systems and data the AI touches, and where every byte lives.
## 2. Change control for AI actions
- How an AI-proposed change is proposed, approved and recorded — the same governed path as any change.
## 3. Guardrails & autonomy ladder
- What the AI may and may not do at each rung; how a rung is earned with evidence.
## 4. Access & least privilege
- Service accounts, scopes, and who holds them.
## 5. Audit & observability
- What is logged, retained, and how an action is reconstructed after the fact.
## 6. Human-in-the-loop
- Where the gates are, who approves, and the stop conditions.
## 7. Sign-off
- CISO / DPO, change board, sponsor — names and date.
Unlocks in module 1.
Structured checkpoint prompts for each of the Cranfield cast, so client feedback is specific and comparable across modules.
# Per-persona feedback templates
Use at each client checkpoint. Match the evidence to what the persona actually fears.
## Dilan Petrou (sponsor, burned by shelfware)
- "Show me this will outlast the engagement." → handover plan + adoption.
- Opening line: the one measurable Week-10 outcome you put your name against.
## Marisol Vega (NOC / change control)
- "Walk me through how an AI-proposed change gets approved." → the human-approval gate.
- She must see: propose-not-apply, evidence shown, an explicit stop.
## Ken Adeyemi (CISO / DPO)
- "Show me the data-access, eval and risk controls in writing." → checkable readiness-gate statements.
- He must see: data residency, least privilege, no unattended production action.
## Priya Raman (service desk, early adopter)
- "Does this genuinely help my team without leaking user data?" → sanitisation + honest deflection metric.
## Tom Böttcher (CFO / procurement)
- "What does this actually save, against where we started?" → measured before/after vs the M2 baseline.
## Renata Cole (your delivery lead)
- "Is this a repeatable land-and-expand play?" → the artifacts + the 90-day expansion plan.
The complete, presentation-ready Cranfield engagement: charter, estate map + baseline, opportunity map, blueprint + readiness gate, delivery foundation, three estate runbooks, integrated runbook pack, and the governance/adoption pack — defended with measured before/after.
The standing control the client keeps: a signable governance pack — data residency, change control for AI actions, guardrails, audit, and the autonomy ladder — written so the CISO and change board actually approve it.
Renata's repeatable play: how one proven Cranfield pilot becomes a path across the estate, sequenced by value and readiness, with the 90-day plan and the artifacts that make it repeatable on the next account.
Go deep on one estate: the DC alert-correlation + capacity + not-suppressed validation runbook, wired to ServiceNow with a human gate, observability and rollback — a working integration, production-ready.
The network root-cause + config-drift + human-gated-change runbook as a working, reversible integration — every AI-proposed change human-approved, with the evidence Marisol would need to trust it.
The retrieval-grounded deflection + bounded self-heal runbook with the sanitisation step and an honest deflection metric, integrated end to end — proving no user data reaches the model.
The strongest capstone is a real engagement you actually face. If your organisation or a client has a live AI-in-operations decision, run the method on that one — the pushback will be free and genuine.