# Forward Deployed Engineer: AI for the Infrastructure Estate

> Embed with a client and land AI into how they run their datacenter, network and workplace — assess, frame, deliver, integrate and govern. 10 modules for infrastructure people who guide and build, not just advise.

- **Audience:** Infrastructure Engineers · Platform & MSP Delivery Leads · Solutions/Field Engineers · Ops Leads embedding AI with clients
- **Level:** Intermediate (infrastructure experience assumed; no ML background required)
- **Duration:** 10 weeks · 1 module/week · 5–7 h/week
- **Modules:** 10
- **Pass mark:** 70%
- **Interactive version:** https://ragentic.netlify.app/#/courses/fde-infrastructure

**This file is generated from the course data by `scripts/build-notes.mjs`. Edit the course data, not this file.**

## The world of this course

**Organisation:** Cranfield Group

Your client for the next ten weeks: Cranfield Group, a mid-market distribution and light-
manufacturing firm — two owned datacenters plus a colo they are half-migrated into, an SD-WAN across 42
sites and 5 plants, and roughly 6,000 endpoints. You are embedded from an infrastructure-support practice
to make AI actually land in how Cranfield runs the estate — not to sell a platform. Dilan, the Head of
Infrastructure, was sold an "AIOps platform" two years ago that became shelfware, and will believe outcomes,
not slides. Your own delivery lead wants a repeatable land-and-expand play out of this. Everything you build
is one accumulating engagement, defended to the room in Week 10.

---

## Phase 1 — Land & Understand (weeks 1–2)

### Module 1 — The Infrastructure FDE: Role, Trust & AI Foundations

**Guiding question:** What is a forward deployed engineer on an infrastructure estate — and how do I earn a sceptical sponsor's trust?

**Outcome:** Define the FDE model and separate it from consulting, pre-sales, and traditional MSP break-fix. Explain forward deployment (embed, iterate, earn trust) versus one-shot tool implementation. Draw the co-pilot vs autopilot line for infrastructure specifically, and establish a shared AI-in-operations vocabulary and the FDE duty of care.

**Estate lens:** You are the person the infrastructure-support firm embeds with a client to make AI actually land in how they run the estate — credible enough to assess it, hands-on enough to build the integration, disciplined enough to be trusted near production. Your deliverable is not a deck; it is a working, governed change that outlasts you. Dilan has been sold "AIOps" before and got shelfware. Trust, earned in weeks, is the whole job.

**Apply-at-work mission — Engagement charter + sealed baseline:** Write the Cranfield engagement charter: your mission, the trust you must earn from Dilan, and what "success by week 10" looks like. Add a stakeholder power/interest map of the cast. Then seal a baseline: in under 150 words, your current unassisted answer to "what can AI realistically do for infrastructure operations, and what can it not?" — Module 10 reopens it.

**Reflection:** Where did my mental model of "AI for infrastructure" turn out to be marketing rather than mechanism? Whose trust on this client will be hardest to earn, and why?

#### Resources

- [LinkedIn Learning — Forward Deployed Engineering in the Age of AI](https://www.linkedin.com/learning/forward-deployed-engineering-in-the-age-of-ai) — Vinoo Ganesh's short course on what forward deployment is and why it differs from SaaS implementation — the mindset this whole programme runs on.
  - Explain forward deployment versus one-shot tool implementation
  - Say why trust, earned on-site, is the FDE's real deliverable
  - Describe how an FDE translates real needs into technical requirements
- [Andrej Karpathy — Intro to Large Language Models](https://www.youtube.com/watch?v=zjkBMFhNj_g) — The one-hour plain-language mental model of what an LLM actually is — enough AI foundation to reason about where it fits an estate, without the maths.
  - Hold a working mental model of what an LLM is and is not
  - Explain fluency-without-truth well enough to set client expectations
  - Ground later AI-in-operations decisions in how the model really works
- [Microsoft Learn — Fundamentals of Generative AI](https://learn.microsoft.com/en-us/training/modules/fundamentals-generative-ai/) — Short, free, enterprise-flavoured grounding in the vocabulary your client's governance and security people will use.
  - Speak the enterprise generative-AI vocabulary stakeholders use
  - Separate what the model does from what the platform wraps around it
  - Place generative AI against the automation the estate already runs
- [Anthropic — Building Effective Agents](https://www.anthropic.com/engineering/building-effective-agents) — The workflow-vs-agent distinction and the co-pilot/autopilot line — from a frontier lab. Read for the judgement, not the code.
  - Make the workflow-vs-agent distinction on a whiteboard
  - Draw the co-pilot vs autopilot line for an infrastructure task
  - Name the guardrail any acting agent needs near production
- [The forward deployed engineer role — curated search](https://www.google.com/search?q=forward+deployed+engineer+role+what+it+is) — Read two or three credible accounts of the FDE role (Palantir-style and modern AI-lab versions). Watch for how it differs from consulting and pre-sales.
  - Contrast the FDE with consulting, pre-sales and MSP break-fix
  - Recognise the failure mode of AI-theatre and shelfware
  - Form your own line on what an infrastructure FDE is accountable for

#### In-world ticket queue

> Day one on the Cranfield account. The kickoff is done; the goodwill is thin. Dilan opened with
> "the last lot sold us a dashboard nobody opened" and gave you ten weeks. Your inbox already has a competing
> appliance quote forwarded by Tom, an eager "can AI just auto-resolve tickets?" from Priya, and a one-line
> note from Ken: "before anything connects to our data, I want to know what and where."

| Ref | Priority | From | Request |
| --- | --- | --- | --- |
| CRN-0001 | P1 | Dilan Petrou | Sponsor kickoff: "prove this won't be shelfware like last time" |
| CRN-0002 | P1 | you, to yourself | Write the engagement charter + who actually decides what |
| CRN-0003 | P1 | Ken Adeyemi | "What data will this touch, and where does it go?" — before any connection |
| CRN-0004 | P2 | Tom Böttcher | Competing appliance quote forwarded — "is this the easy button?" |

#### Project — Engagement charter + sealed baseline (Capstone component)

Write the Cranfield engagement charter: your mission, the specific trust you must earn from Dilan (who was burned by shelfware), and what "success by week 10" looks like as observable outcomes. Add a power/interest stakeholder map of the six-person cast — who decides, who can block, who to win first. Then seal a baseline: in under 150 words, your current unassisted answer to "what can AI realistically do for infrastructure operations, and what can it not?" Date it and set it aside — Module 10 reopens it.

**Deliverable:** fde/m01-charter-and-baseline.md — the charter, the stakeholder map, and the dated sealed baseline. This is the first component of your accumulating engagement pack.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Outcome-framed charter | 30% | Success is stated as observable Week-10 outcomes Dilan would accept, not activities or deliverable counts. |
| Honest stakeholder map | 30% | Each of the six cast members placed on power/interest with a specific reason; names who to win first and why. |
| Genuine sealed baseline | 25% | An unassisted, dated explanation written before the course changes it — its errors are the point, not polish. |
| FDE framing | 15% | Frames the engagement as a governed change that outlasts you, not a demo or a report. |

#### Scenario drills

**Drill 1.** Kickoff. Dilan opens with "the last lot sold us a dashboard nobody opened" (CRN-0001) and gives you ten weeks.

**Task:** Write the two-sentence reply that neither dismisses the past failure nor over-promises, and names the one thing you'll do differently. Then state the single Week-10 outcome you'd put your name against.

**Drill 2.** Ken (CRN-0003): "Before anything connects to our data, I want to know what and where."

**Task:** List the three questions you must answer for Ken before any tool touches Cranfield data, and say why each matters more than the demo he hasn't asked for.

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.**

**1. The clearest difference between a forward deployed engineer and a traditional consultant is that the FDE…**

   a. Charges the client a considerably higher day rate than the consultant does
   b. Embeds to build and land the change
   c. Works remotely from the head office
   d. Writes the strategy slide deck

**2. Forward deployment differs from a normal SaaS/tool implementation mainly because it…**

   a. It simply installs the software faster
   b. It iterates with the client to fit real needs and earn their trust over time
   c. It needs no product configuration
   d. It always costs less than a rollout

**3. For an infrastructure FDE, the "co-pilot vs autopilot" line is about…**

   a. Whether the underlying tool happens to have a chat interface bolted onto it
   b. Whether AI proposes for a human or acts alone
   c. Which cloud region the model happens to be hosted in
   d. How quickly the model returns its answer to you

**4. Dilan was "burned by shelfware" before. The main lesson for your engagement is that…**

   a. You should quietly avoid ever mentioning the previous failed vendor to him
   b. Trust is earned through measured outcomes
   c. He approves anything to avoid conflict
   d. The old tool just needs reinstalling

**5. An LLM sounds fluent and confident even when wrong because…**

   a. It is deliberately trying to mislead whoever happens to be asking it
   b. Fluent plausible text is the goal, not truth
   c. Only the cheaper, older models tend to do this
   d. It briefly lost its connection

**6. "AI theatre" on an infrastructure account most often looks like…**

   a. A carefully governed pilot
   b. An impressive demo that solves no real, owned problem and quietly becomes shelfware
   c. A boring automation that saves hours
   d. A tightly scoped, measured proof

**7. The FDE's "duty of care" on a client estate primarily means…**

   a. Making sure you always adopt the very newest model available on the market
   b. Not letting AI create risk the client cannot see
   c. Keeping your billed day rate as low as it can possibly go
   d. Putting absolutely everything into one very long report

**8. The strongest first move with a sceptical sponsor like Dilan is to…**

   a. Promise a dramatic, estate-wide transformation within the first few weeks
   b. Agree what a measurable success looks like
   c. Build the flashiest pilot you can immediately
   d. Ask him to trust the process and wait

**9. A stakeholder power/interest map is useful because it…**

   a. It quietly replaces the need to actually go and talk to any of the people
   b. Shows who can block you and who to win first
   c. Is mandatory for every IT project
   d. Measures how technically skilled each person is

**10. Which is genuinely NOT part of the infrastructure FDE role as this course frames it?**

   a. Assessing the client's estate
   b. Training a foundation model from scratch
   c. Guiding the client and building the integration work
   d. Governing the AI once it runs

**11. The sealed baseline you write in Module 1 exists so that…**

   a. The course simply needs a week-one assignment for the tutor to grade
   b. You can honestly measure how your thinking changed
   c. Your manager simply needs a document to file for compliance
   d. You can reuse it word-for-word as your final answer

**12. Treating the engagement as "a governed change that outlasts you" mainly implies…**

   a. You should stay embedded with them forever
   b. The client's own team can keep running it long after you have left
   c. The solution must use the largest model
   d. You avoid writing anything down

### Module 2 — Reading the Estate: Assessment & Discovery

**Guiding question:** Before I recommend anything, what is actually in this estate, where does it hurt, and where does usable data already live?

**Outcome:** Run a structured assessment across datacenter, network and workplace. Inventory the data sources that matter — telemetry, event/alert streams, tickets, logs, configs, topology, CMDB — and rate each for AI-readiness. Establish a measured baseline (MTTR, alert volume, ticket mix, capacity headroom) you can later prove improvement against.

**Estate lens:** Discovery for infrastructure is not a workshop of opinions; it is reading the estate's own exhaust. The CMDB is ~70% true, the monitoring is a three-tool patchwork, and the pain everyone complains about is not always where the data is. Find both. The baseline you take this week is the honest yardstick your Module-10 outcomes get measured against — take it properly or the whole engagement's "impact" is a feeling.

**Apply-at-work mission — Estate map, data inventory & baseline:** Build the Cranfield estate + data inventory: what telemetry, tickets, configs and CMDB data exist, the trust level of each, and where the CMDB lies. Record a measured baseline for the pains you will target. Mark every unknown with a "?" rather than a guess — the gaps are the engagement. 🎯 Completes Capstone Milestone 1.

**Reflection:** Which "obvious" pain point turned out to have no usable data behind it? Which data source surprised me by being richer than expected?

#### Resources

- [Atlassian — ITSM incident & request management](https://www.atlassian.com/itsm) — A clean refresher on the ticket/incident/request lifecycle you'll read Cranfield's ServiceNow data against during discovery.
  - Recognise the incident/request/change data an estate generates
  - Know what a healthy ITSM record contains before you mine it
  - Map where AI could help each lifecycle stage — and where it can't
- [Configuration management & the CMDB — curated search](https://www.google.com/search?q=CMDB+configuration+management+database+accuracy+problems) — Read one credible piece on why CMDBs drift out of date. Cranfield's is ~70% trustworthy — assume yours is too until proven otherwise.
  - Explain why a CMDB is rarely fully trustworthy
  - Plan to reconcile CMDB claims against live telemetry
  - Treat inventory as a data-quality problem, not a given
- [Google SRE Book — Monitoring distributed systems](https://sre.google/sre-book/monitoring-distributed-systems/) — The four golden signals and the difference between symptoms and causes — the vocabulary for reading a monitoring patchwork like Cranfield's.
  - Use symptom-vs-cause and the golden signals when reading telemetry
  - Judge whether an alert stream is signal or noise
  - Identify which monitoring data is AI-ready and which is junk
- [UK ICO — anonymisation and pseudonymisation guidance](https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/) — Estate telemetry, tickets and logs are identity-dense (hostnames, users, IPs). The formal grounding for the sanitisation discipline you'll need before any of it reaches a model.
  - Distinguish anonymisation from pseudonymisation for estate data
  - Spot the identity payload in telemetry, tickets and logs
  - Justify a sanitisation rule to a DPO in their own terms

#### In-world ticket queue

> Discovery week. The CMDB export arrived and roughly a third of it is wrong. Monitoring is three
> tools that disagree. Everyone "knows" what hurts most, and no two people agree. You need the estate's own
> data, not its opinions — and a baseline you can be measured against in Week 10.

| Ref | Priority | From | Request |
| --- | --- | --- | --- |
| CRN-0011 | P1 | you, to yourself | CMDB export is ~70% trustworthy — reconcile against reality |
| CRN-0012 | P2 | Marisol Vega | Which of SolarWinds / Datadog / Zabbix is the source of truth? |
| CRN-0013 | P1 | you, to yourself | Baseline the pain: MTTR, alert volume, ticket mix, capacity headroom |
| CRN-0014 | P2 | Ken Adeyemi | "Where does the telemetry and ticket data actually live?" |

#### Project — Estate map, data inventory & baseline (Capstone component)

Build Cranfield's current-state estate map across datacenter, network and workplace, and a data inventory: for each source (telemetry, event/alert streams, tickets, logs, configs, CMDB) record what it is, how trustworthy it is, and rate it for AI-readiness. Then record a measured baseline for the pains you expect to target — MTTR, alert volume, ticket mix, capacity headroom — the honest yardstick Module 10 measures against. Mark every unknown with a "?" rather than a guess.

**Deliverable:** fde/m02-estate-map-and-baseline.md — the estate map, the rated data inventory, and the dated baseline with question marks preserved. Second component of the engagement pack.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Cross-estate coverage | 25% | Datacenter, network and workplace all mapped, with the systems of record (ServiceNow, the monitoring patchwork, CMDB) located. |
| Honest data-readiness rating | 30% | Each source rated for trust and AI-readiness with a reason; the CMDB's unreliability is confronted, not assumed away. |
| Measured baseline | 30% | Concrete before-numbers captured for the targeted pains, dated, so Week-10 improvement can be proven honestly. |
| Question marks preserved | 15% | Unknown cells marked unknown rather than plausibly filled — the gaps are the discovery output. |

#### Scenario drills

**Drill 1.** The CMDB export lands and ~30% is wrong (CRN-0011): decommissioned hosts present, live services missing.

**Task:** Describe how you'd reconcile it against something more trustworthy in a week, and name the one class of error you'd check first before trusting any of it for a pilot.

**Drill 2.** Three monitoring tools disagree on whether DC-1 is healthy (CRN-0012); Marisol asks which to believe.

**Task:** Give the two-line rule you'd use to decide the source of truth per signal, and one question that would settle it for the specific case of DC-1 capacity.

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.**

**1. Discovery on an infrastructure estate should rely most on…**

   a. A workshop where stakeholders get together and share their opinions about what hurts
   b. The estate's own telemetry, tickets and configs
   c. The vendor's recommended reference architecture diagram
   d. The single most senior person's gut intuition about the estate

**2. Cranfield's CMDB is about 70% trustworthy. The right response is to…**

   a. Treat it as authoritative
   b. Reconcile every load-bearing claim in it against live telemetry and the ticket record
   c. Discard it entirely as useless
   d. Ask the client to rebuild it first

**3. The point of taking a measured baseline in Module 2 is to…**

   a. Fill in a required field on the project template so the pack looks complete
   b. Let Week-10 prove improvement honestly
   c. Give the CFO a number to approve the budget right now today
   d. Compare against benchmarks

**4. The "four golden signals" of monitoring are most useful to an FDE for…**

   a. Choosing which single monitoring vendor the client should standardise everything onto
   b. Judging whether an alert stream is signal or noise
   c. Costing the observability stack
   d. Deciding NOC headcount

**5. Before any Cranfield telemetry or ticket data reaches a model, you must…**

   a. Get written sign-off from the entire board before anything at all can happen
   b. Strip identity: hostnames, users, IPs
   c. Convert everything into one single spreadsheet first
   d. Wait until the monitoring tools are fully consolidated

**6. A data source that is high-volume but low-trust (like a noisy alert feed) should be…**

   a. Used immediately
   b. Flagged low AI-readiness until its signal is actually proven out on real data
   c. Deleted from the inventory
   d. Treated like a clean source

**7. The most useful thing to do with a data source you cannot yet assess is to…**

   a. Assume it is basically fine and just move quickly on to the next source
   b. Mark it as unknown rather than guessing
   c. Leave it out of the inventory entirely for now
   d. Rate it high so things look complete

**8. Symptom-versus-cause matters in discovery because the loudest pain…**

   a. Is always the single most valuable thing you could possibly automate first
   b. May be a symptom whose data lives elsewhere
   c. Should just be ignored
   d. Is caused by the tool

**9. Cranfield runs SolarWinds, Datadog and Zabbix that disagree. For discovery this means…**

   a. You must pick just one of them and delete the other two immediately
   b. Reconcile them to find the truth
   c. The client is clearly incompetent at monitoring their estate
   d. Monitoring data cannot be used for AI

**10. The estate map is most valuable to the engagement as…**

   a. A one-off diagram
   b. A living artifact that every later module in the engagement builds on and enriches
   c. Evidence you did the work
   d. A doc only you read

**11. The biggest risk of skipping a proper baseline is that…**

   a. The overall project plan ends up being very slightly shorter than it was
   b. You cannot honestly prove impact
   c. The client asks for a bit more documentation later
   d. The tools generate more alerts

**12. When Ken asks "where does the telemetry and ticket data actually live?", the honest FDE answer is…**

   a. "Wherever the particular AI tool happens to need it to physically be at the time"
   b. A mapped account of each source and where it lives
   c. "That's the vendor's problem"
   d. "We'll figure it out later"

---

## Phase 2 — Find & Frame (weeks 3–4)

### Module 3 — Use-Case Mining, Value & Risk Triage

**Guiding question:** Across the estate, where does AI genuinely earn its place — and what is each opportunity worth against its risk?

**Outcome:** Build the infrastructure-AI opportunity catalogue (AIOps and event correlation, observability and anomaly, capacity forecasting, config and change automation, ticket deflection and knowledge, security operations, DC energy). Score each on value × data-readiness × risk, set a quantified KPI hypothesis for the shortlist, and make the buy-vs-integrate-vs-build call. Kill the solution-first ideas early.

**Estate lens:** The fastest way to become the next shelfware is to demo a clever AI thing nobody needed. The FDE's core judgement is triage: which of the estate's real pains has both high value and enough usable data to be provable, and should you buy a tool, integrate one, or build. A quantified KPI hypothesis — "cut DC-1 alert noise 40% without hiding a real incident" — is what turns a good idea into something you can be held to.

**Apply-at-work mission — Opportunity map + KPI hypotheses:** Score eight candidate opportunities for Cranfield on value × data-readiness × risk. Shortlist three, and for each state a quantified KPI hypothesis measured against your Module-2 baseline, a risk tier, and a buy/integrate/build recommendation with rationale.

**Reflection:** Which opportunity was I tempted by because it was interesting rather than valuable? Which "boring" one scored highest, and why?

#### Resources

- [IBM — What is AIOps?](https://www.ibm.com/topics/aiops) — A vendor-neutral grounding in the AIOps opportunity space — event correlation, anomaly, noise reduction — so you can name candidates for Cranfield without buying the hype.
  - Name the main AIOps opportunity categories for an estate
  - Separate a real correlation use case from a dashboard
  - Judge which AIOps claims need data you may not have
- [PagerDuty — What is alert fatigue?](https://www.pagerduty.com/resources/learn/what-is-alert-fatigue/) — Why DC-1's 900-alerts-a-night problem is worth money to fix — and why "fewer alerts" is a dangerous goal if you cannot prove nothing real was hidden.
  - Quantify the operational cost of alert fatigue
  - Frame noise reduction as a value opportunity with a risk
  - Set a KPI that rewards signal, not just silence
- [ProductPlan — The RICE scoring model](https://www.productplan.com/glossary/rice-scoring-model/) — A clean prioritisation frame you can adapt to value × data-readiness × risk. Borrow the discipline of scoring, not the exact factors.
  - Adapt a scoring model to value, data-readiness and risk
  - Force a ranked shortlist instead of a wish list
  - Defend why one opportunity beats another with numbers
- [Build vs buy vs integrate enterprise AI — curated search](https://www.google.com/search?q=build+vs+buy+vs+integrate+enterprise+AI+decision) — Read two credible takes on the buy/integrate/build decision for enterprise AI. For an FDE on an estate, integrate is usually the default and build the exception.
  - State when to buy, integrate, or build an AI capability
  - Recognise why "build" is rarely the right first move
  - Tie the decision to data ownership and time-to-value
- [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework) — The reference for tiering AI risk. Use it to justify why a security-ops or auto-change idea carries more risk than ticket deflection.
  - Assign a defensible risk tier to an AI opportunity
  - Explain why some estate use cases need heavier controls
  - Speak the risk vocabulary a CISO like Ken will accept

#### In-world ticket queue

> Framing week. Now the pet ideas surface: Priya wants auto-resolve, a plant manager wants "predictive
> everything", and the vendor swears their appliance does it all. Tom wants to know which of these is worth
> money. Your job is triage — value against risk against whether the data even exists.

| Ref | Priority | From | Request |
| --- | --- | --- | --- |
| CRN-0021 | P1 | you, to yourself | Rank the AI opportunities across DC / network / workplace |
| CRN-0022 | P3 | Operations | "Predictive everything" request from the plant — real or hype? |
| CRN-0023 | P1 | Tom Böttcher | Tom: "which two of these actually save money, and how much?" |
| CRN-0024 | P2 | Priya Raman | Priya: "can we at least start with password/VPN deflection?" |

#### Project — Opportunity map + KPI hypotheses (Capstone component)

Score eight candidate AI opportunities for Cranfield across datacenter, network and workplace on value × data-readiness × risk, using the Module-2 inventory as the readiness input. Shortlist three. For each of the three, state a quantified KPI hypothesis measured against your Module-2 baseline (e.g. "cut DC-1 alert noise 40% without suppressing a real incident"), assign a risk tier, and make a buy / integrate / build recommendation with a one-line rationale. Explicitly name one solution-first idea you killed and why.

**Deliverable:** fde/m03-opportunity-map.md — the scored catalogue, the shortlist of three with KPI hypotheses and risk tiers, the buy/integrate/build calls, and the killed idea. Third component of the engagement pack.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Honest scoring | 30% | Each opportunity scored on value, data-readiness and risk with a stated reason; readiness ties back to the Module-2 inventory, not optimism. |
| Quantified KPI hypotheses | 30% | Each shortlisted item has a specific, measurable target tied to the Module-2 baseline — a number you could be held to, not a vibe. |
| Buy/integrate/build judgement | 25% | A defensible recommendation per shortlisted item, with integrate as the reasoned default and build justified only where nothing fits. |
| A killed idea | 15% | Names a genuinely tempting solution-first idea and explains why value or data killed it — evidence of triage, not enthusiasm. |

#### Scenario drills

**Drill 1.** Tom forwards a competing appliance quote and asks "which two of these actually save money, and how much?" (CRN-0023).

**Task:** Pick the two Cranfield opportunities you would put in front of Tom, and for each give the one number you would commit to against the Module-2 baseline. Say what you would NOT promise.

**Drill 2.** Priya asks "can we at least start with password/VPN deflection?" (CRN-0024) while a plant wants "predictive everything".

**Task:** Rank these two on value × data-readiness × risk in three sentences, and say which you shortlist first and why.

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.**

**1. The core triage of value × data-readiness × risk exists mainly to…**

   a. Make the opportunity catalogue look more thorough than it really is to the client
   b. Stop you building things nobody can prove
   c. Ensure the newest available model gets used somewhere
   d. Give the procurement team a longer shortlist to review

**2. The fastest route to becoming the next shelfware is to…**

   a. Ship a boring automation
   b. Demo a clever AI thing that nobody actually needed and that no one will own
   c. Score ideas against the baseline
   d. Recommend integrating a tool

**3. A quantified KPI hypothesis matters because it…**

   a. Fills in an otherwise empty field on the opportunity-scoring template
   b. Turns an idea into a measurable target
   c. Reliably impresses the finance team in the room
   d. Shortens the plan

**4. The plant's "predictive everything" request is best treated as…**

   a. A top priority the team should just start building on right away immediately
   b. Hype until the data to predict on is shown to exist
   c. Proof the client gets AI
   d. A reason to buy the appliance

**5. You would lean toward BUYING a tool when…**

   a. The requirement is genuinely unusual and no product exists for it at all
   b. A mature product fits and integrates cheaply
   c. You want to keep maximum control over all of the code yourself
   d. The data is far too sensitive to ever be moved anywhere

**6. You would consider BUILDING (the rare case) only when…**

   a. A product nearly fits and could be configured
   b. The need is specific and valuable and genuinely no product or integration fits it
   c. The vendor demo looked impressive
   d. You have spare time in the plan

**7. A high-value opportunity with no usable data behind it should be…**

   a. Parked until the data to support it exists
   b. Shortlisted anyway because the value looks high
   c. Built using synthetic data as a starting point
   d. Sold to the CFO

**8. Ticket deflection scored high for Cranfield mainly because it is…**

   a. It is simply the newest and most fashionable thing on the whole opportunity list
   b. High-volume, repetitive, and rich in data from Priya's runbooks
   c. It is low-risk only
   d. It was the vendor's idea

**9. Security-operations AI warrants a heavier risk tier because…**

   a. It happens to be a particularly fashionable area right now among vendors
   b. False negatives are costly; data is sensitive
   c. It usually needs a considerably bigger model to work
   d. The vendor strongly recommends treating it that way

**10. Scoring opportunities against the Module-2 baseline matters because…**

   a. The scoring template asks for it
   b. A value claim only means something when it is tied to a real dated before-number
   c. It fills the slide
   d. It flatters the demo

**11. The right output of the triage step is…**

   a. A single pet idea to take straight into build without further comparison
   b. A ranked shortlist with a reason for each pick
   c. Every idea kept in the catalogue for possible use later on
   d. A vendor quote

**12. Killing a solution-first idea early is best read as a sign of…**

   a. A worrying kind of indecision that clients will quickly come to distrust
   b. FDE judgement that protects the client's trust and budget
   c. A simple lack of ambition
   d. A quiet fear of building

### Module 4 — Solution Blueprint & the Readiness Gate

**Guiding question:** How do I turn a chosen opportunity into a blueprint I can build safely — and prove I am actually ready to start?

**Outcome:** Produce a solution blueprint: architecture, scope, dependencies, delivery plan, and acceptance criteria. Reason about blast radius and the advisory-before-action ladder for anything touching production. Then pass the readiness gate — explicit, checkable statements of data/access assumptions, evaluation criteria, and risk controls, without which implementation does not begin.

**Estate lens:** This module is the bridge from consulting conversation to engineering. A learner can understand a client's problem perfectly and still fail in delivery if they never converted it into a blueprint with data contracts, acceptance criteria, and controls. The readiness gate is a real stage-gate, not a formality: if you cannot write your data-access assumptions, eval criteria and risk controls as checkable statements, you are not ready to build — and saying so is the senior move.

**Apply-at-work mission — Blueprint + checkable readiness gate:** Produce the Cranfield solution blueprint (architecture, scope, dependencies, delivery plan, acceptance criteria) for your top pilot. Then write the readiness gate as checkable statements: data/access assumptions, evaluation criteria, risk controls, and the human-approval path. A colleague should be able to tick each one true or false. 🎯 Completes Capstone Milestone 2.

**Reflection:** Which assumption did writing it down as a checkable statement expose as wishful? What would break first if I skipped the gate and started building?

#### Resources

- [AWS — Well-Architected Framework](https://aws.amazon.com/architecture/well-architected/) — The reference for reasoning about a solution across reliability, security, and operations — adapt its discipline to your Cranfield pilot blueprint.
  - Structure a blueprint across reliability, security and ops
  - Surface dependencies and failure points before building
  - Justify architecture choices a client can review
- [Atlassian — Acceptance criteria](https://www.atlassian.com/agile/project-management/user-stories) — How to write acceptance criteria that make "done" checkable. The readiness gate lives or dies on statements a colleague can tick true or false.
  - Write acceptance criteria as checkable statements
  - Turn a vague goal into a testable definition of done
  - Separate what must be true from what would be nice
- [Google SRE Workbook — Canarying releases](https://sre.google/workbook/canarying-releases/) — The discipline of limiting blast radius when rolling out change — directly applicable to how an AI-proposed action should reach production.
  - Reason about blast radius for an AI-driven change
  - Design a staged, reversible path to production
  - Decide what to canary before trusting an action
- [Evaluating LLM applications — curated search](https://www.google.com/search?q=how+to+evaluate+LLM+applications+evals+acceptance+criteria) — Read two credible pieces on building evals for an LLM feature. You must be able to state eval criteria as part of the readiness gate before any build begins.
  - Define eval criteria for an AI feature before building it
  - Choose measures that reflect real acceptance, not vibes
  - Make evals part of the go/no-go, not an afterthought
- [NIST AI RMF — Playbook (Govern & Map)](https://airc.nist.gov/AI_RMF_Knowledge_Base/Playbook) — Concrete control language for the risk-controls half of your readiness gate — the words Ken will expect to see in writing before anything connects.
  - Express risk controls as concrete, checkable statements
  - Map data-access assumptions to least-privilege controls
  - Give a CISO the written controls before connecting anything

#### In-world ticket queue

> Blueprint week. Dilan has approved a shortlist in principle, but Ken will not let anything start
> without controls, and Marisol wants to see exactly how a change gets proposed and approved. You cannot enter
> build until the assumptions, evals and risk controls are written down and checkable.

| Ref | Priority | From | Request |
| --- | --- | --- | --- |
| CRN-0031 | P1 | you, to yourself | Solution blueprint + delivery plan for the top pilot |
| CRN-0032 | P1 | Ken Adeyemi | Ken: "show me the data-access, eval and risk controls in writing" |
| CRN-0033 | P1 | Marisol Vega | Marisol: "walk me through how an AI-proposed change gets approved" |
| CRN-0034 | P2 | you, to yourself | Readiness gate: can a colleague tick every assumption true or false? |

#### Project — Blueprint + checkable readiness gate (Capstone component)

Produce the Cranfield solution blueprint for your top shortlisted pilot: architecture, scope, dependencies, delivery plan, and acceptance criteria. Reason explicitly about blast radius and place the pilot on the advisory-before-action ladder. Then write the readiness gate as checkable statements grouped into data/access assumptions, evaluation criteria, risk controls, and the human-approval path — each phrased so a colleague can tick it true or false. If any statement cannot yet be ticked true, say so: that is the gate doing its job.

**Deliverable:** fde/m04-blueprint-and-readiness-gate.md — the blueprint and the readiness gate as checkable statements, with any un-tickable items flagged. Fourth component of the engagement pack; completes Capstone Milestone 2.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Buildable blueprint | 30% | Architecture, scope, dependencies and delivery plan are concrete enough to build from; acceptance criteria state what "done and good" means. |
| Blast-radius reasoning | 25% | Places the pilot honestly on the advisory-before-action ladder and reasons about what a wrong action could touch. |
| Checkable readiness gate | 30% | Data/access, eval and risk-control statements are written so a colleague can tick each true or false — not aspirational prose. |
| Honest not-ready flags | 15% | Any assumption that cannot yet be ticked true is flagged rather than hidden — the senior move the gate exists to reward. |

#### Scenario drills

**Drill 1.** Ken: "show me the data-access, eval and risk controls in writing" before anything starts (CRN-0032).

**Task:** Draft three checkable readiness-gate statements — one data-access, one eval, one risk control — each phrased so Ken can tick it true or false. Flag any you cannot yet tick true.

**Drill 2.** Marisol: "walk me through how an AI-proposed change gets approved" (CRN-0033).

**Task:** Describe the advisory-before-action path for one Cranfield change: what the AI proposes, what it must show, who approves, and where the hard stop is.

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.**

**1. This module is the bridge from consulting conversation to…**

   a. A noticeably longer and more detailed slide deck for the sponsor
   b. Engineering you can build from
   c. A formal vendor selection exercise
   d. A recurring weekly status meeting

**2. The readiness gate is best understood as…**

   a. A formality to sign
   b. A real stage-gate: no build begins until every assumption is checkable
   c. A document for the client
   d. A vendor requirement

**3. Acceptance criteria are useful mainly because they…**

   a. They are strictly required by the template process the team happens to follow
   b. Make "done" checkable rather than opinion
   c. Reliably impress the sponsor in the review meeting
   d. Fill the plan

**4. Reasoning about blast radius before building means asking…**

   a. How recent and how large the underlying model the team has chosen actually is
   b. What a wrong action could touch, and how far damage spreads
   c. Which vendor is cheapest
   d. How fast it runs

**5. The "advisory-before-action" ladder says that near production you should…**

   a. Automate the action first and then add human oversight later if it turns out to be needed
   b. Start with AI advising, and earn each rung
   c. Let the agent act and simply log what it did afterwards
   d. Skip straight to full autonomy for the sake of speed

**6. A data-access assumption belongs in the readiness gate as…**

   a. A hope
   b. A checkable statement a colleague can mark true or false against reality
   c. A vendor promise
   d. A verbal agreement

**7. Eval criteria should be defined…**

   a. Only after the pilot has been fully built, to see how well it happened to do
   b. Before the build, as part of the go/no-go
   c. Only if the client specifically asks for them
   d. By the vendor

**8. Writing an assumption down as a checkable statement most often…**

   a. It largely wastes time that would be far better spent actually building the thing
   b. Exposes which "facts" were actually wishful thinking
   c. It slows the client down
   d. It duplicates the blueprint

**9. The human-approval path in the gate specifies…**

   a. Who exactly should be assigned the blame if the whole thing goes badly wrong
   b. Who approves an action, and when
   c. The main vendor contact for the project
   d. The order the demo should be given in

**10. If you cannot write your risk controls as checkable statements, the right move is to…**

   a. Start building and refine later
   b. Say you are not ready — declaring that honestly is the senior move here
   c. Ask the vendor to write them
   d. Skip the controls entirely

**11. Naming dependencies in the blueprint matters because…**

   a. It usefully lengthens the overall document and makes it look more rigorous to the client
   b. An unnamed dependency is what blocks the build
   c. Clients simply expect to be given a fairly long list
   d. It looks rigorous

**12. Scope in the blueprint is best kept…**

   a. Kept as broad as it possibly can be, so as to properly show the client real ambition
   b. Bounded to one provable pilot, the rest as later phases
   c. Left undefined until the client decides
   d. Whatever the vendor scopes

---

## Phase 3 — Prepare the Ground (weeks 5)

### Module 5 — Cross-Estate AI Delivery Foundation

**Guiding question:** What delivery primitives must be in place before I touch any estate — and which estate will I go deepest on?

**Outcome:** Stand up the reusable machinery every estate lab depends on. Block A (data & access): data sources and contracts, ITSM/CMDB and API access, identity/RBAC and least-privilege service accounts, ticket/event ingestion, runbook anatomy. Block B (safety & control): tool/function calling, evaluations and acceptance criteria, guardrails, observability, human-in-the-loop gates, fallback paths, and common failure modes. Then nominate your primary estate for deeper work.

**Estate lens:** This is the module the earlier design was missing. Infrastructure people are strong on estates and weak on agentic patterns and evaluation — so if runbook orchestration, tool calling, guardrails, evals and observability first appear only inside the estate labs, the labs collapse under the load. Learn the delivery primitives once, here, as reusable scaffolding; then the estate modules apply them rather than teach them. This is also where you choose the estate you will go deepest on.

**Apply-at-work mission — Delivery foundation + primary-estate choice:** Produce the Cranfield data/integration/eval/guardrail plan and stand up the shared scaffold: one permissioned tool/API connection, one grounded assistant, one guarded tool call with an eval and a human-in-the-loop gate. Start the production-readiness checklist you will update every later module. Finally, nominate your PRIMARY estate (Datacenter, Network or Workplace) — it gets your deeper lab and more of the Module-9 integration weight.

**Reflection:** Which delivery primitive did I underestimate — access/RBAC, evals, guardrails, observability, or fallbacks? Why did I pick the primary estate I chose?

#### Resources

- [Anthropic — Tool use (function calling)](https://docs.anthropic.com/en/docs/build-with-claude/tool-use/overview) — How a model takes a bounded, permissioned action through a defined interface — the primitive under every agentic runbook you will build in the estate labs. Read for the pattern, not the SDK.
  - Explain how tool/function calling lets a model act safely
  - Describe the interface contract a tool call needs
  - Spot where a tool call must be gated near production
- [Hamel Husain — Your AI product needs evals](https://hamel.dev/blog/posts/evals/) — The practitioner's case for building evaluations before you trust an AI feature — the acceptance-criteria half of your readiness gate, made concrete.
  - Design a repeatable eval for an AI ops feature
  - Tie eval pass/fail to acceptance criteria, not vibes
  - Use evals as the go/no-go before and after building
- [OWASP — Top 10 for LLM Applications](https://owasp.org/www-project-top-10-for-large-language-model-applications/) — The canonical catalogue of LLM failure modes and attacks — prompt injection, data leakage, excessive agency. The backbone of the guardrails and failure-modes toolkit.
  - Name the LLM failure modes that matter on an estate
  - Map each risk to a concrete guardrail
  - Justify guardrail choices in OWASP terms to a CISO
- [Microsoft — Azure RBAC best practices](https://learn.microsoft.com/en-us/azure/role-based-access-control/best-practices) — Least-privilege, scoped access and service-account discipline — the controls Ken will require in writing before any AI connects to Cranfield's systems.
  - Scope a least-privilege service account for an AI task
  - Explain read-only vs write access in RBAC terms
  - Give a CISO the access controls before connecting
- [OpenTelemetry — Observability primer](https://opentelemetry.io/docs/concepts/observability-primer/) — The grounding for making an AI workflow observable — traces, metrics, logs — so you can see what the AI did, why, and whether it still behaves.
  - Define what observability means for an AI workflow
  - Choose signals that reveal AI misbehaviour early
  - Design dashboards the ops team can actually run

#### In-world ticket queue

> Prepare-the-ground week. Before touching any estate you need the plumbing: read-only API access to
> ServiceNow and monitoring, least-privilege service accounts Ken will actually sign, an eval harness, guardrails,
> observability, and a human-in-the-loop gate. Decide which estate you go deepest on before the labs begin.

| Ref | Priority | From | Request |
| --- | --- | --- | --- |
| CRN-0041 | P1 | Ken Adeyemi | Request least-privilege service accounts + read-only API scopes |
| CRN-0042 | P1 | you, to yourself | Stand up the eval + guardrail + HITL scaffold once, reuse everywhere |
| CRN-0043 | P2 | Renata Cole | Start the production-readiness checklist (living document) |
| CRN-0044 | P2 | you, to yourself | Nominate your PRIMARY estate: Datacenter, Network or Workplace |

#### Project — Delivery foundation + primary-estate choice (Capstone component)

Produce Cranfield's data / integration / eval / guardrail plan, then stand up the shared scaffold the estate labs will reuse: one permissioned tool or API connection (read-only, least-privilege), one grounded assistant over real Cranfield content, and one guarded tool call with an eval and a human-in-the-loop gate. Start the production-readiness checklist you will update every later module. Finally, nominate your PRIMARY estate — Datacenter, Network or Workplace — with a one-paragraph rationale; it gets your deeper lab and more of the Module-9 integration weight.

**Deliverable:** fde/m05-delivery-foundation.md — the data/access/eval/guardrail plan, the described scaffold (tool connection, grounded assistant, guarded call + eval + HITL gate), the started production-readiness checklist, and your primary-estate choice. Fifth component of the engagement pack.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Data & access plan | 25% | Data sources, contracts, least-privilege service accounts and read-only API scopes are specified concretely enough for Ken to sign. |
| Safety & control scaffold | 30% | A guarded tool call with an eval and an explicit human-in-the-loop gate is designed — the reusable machinery, not a one-off. |
| Production-readiness checklist started | 20% | A living checklist is begun with real items (evals, rollback, observability, HITL, change record) to be updated each later module. |
| Reasoned primary-estate choice | 25% | Names a primary estate with a rationale tied to value, data-readiness and the learner's own development — not an arbitrary pick. |

#### Scenario drills

**Drill 1.** Ken will sign least-privilege service accounts and read-only API scopes — but wants them specified first (CRN-0041).

**Task:** Write the access request for one Cranfield integration: which system, read or write, scoped to what, and why that scope and no wider. Name the one thing you are deliberately NOT asking for.

**Drill 2.** You must stand up the shared scaffold once and reuse it across all three estates (CRN-0042).

**Task:** List the four reusable primitives you would build first and, for one guarded tool call, describe the eval and the human-in-the-loop gate you would wrap around it.

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.**

**1. Why learn the delivery primitives here in Module 5 rather than inside each estate lab?**

   a. Because the client specifically asked for a dedicated foundation week up front
   b. Learn them once so the labs apply, not teach
   c. Because the estate labs would otherwise collapse under the combined load
   d. Because the vendor tooling requires it before anything else begins

**2. A least-privilege service account is…**

   a. A shared admin login
   b. An account scoped to only the exact read and write access the task genuinely needs
   c. The vendor default account
   d. A personal named account

**3. A data contract mainly defines…**

   a. A binding legal agreement the client signs with the AI model vendor before any use
   b. What a source provides and in what shape
   c. The commercial pricing terms for the data feed being consumed
   d. The uptime SLA

**4. An evaluation ("eval") in this context is…**

   a. A one-off subjective impression of how good the demo happened to feel on the day
   b. A repeatable test of whether the AI meets its acceptance criteria
   c. A vendor benchmark
   d. A user survey

**5. A guardrail is best described as…**

   a. A dashboard that displays the AI system's activity to the operations team
   b. A limit on what the AI may do
   c. A model that has been fine-tuned on the client's own private data
   d. A written policy document filed with the compliance team

**6. Observability for an AI workflow means…**

   a. Making it look good
   b. Being able to see what the AI did, why, and whether it is still behaving as intended over time
   c. A monthly report
   d. A public status page

**7. A human-in-the-loop gate is…**

   a. A mandatory training session every operator must attend before using the system at all
   b. A point where a human must approve before the action proceeds
   c. A recurring weekly committee that reviews all completed actions
   d. A login screen

**8. A fallback path matters because…**

   a. Because the client will always insist on having some kind of manual backup anyway
   b. What the system does when the AI is unsure or fails is itself a safety decision
   c. It looks thorough
   d. The vendor includes one

**9. Learning runbook anatomy is worth it because…**

   a. Because runbooks are a legacy artifact the client will eventually want to retire fully
   b. An AI runbook automates a structure ops already trusts
   c. Because the compliance team maintains a strict template for exactly how they must be written
   d. Because every monitoring vendor ships its own competing runbook format you must learn

**10. Tool/function calling lets a model…**

   a. Write poetry
   b. Take a bounded, permissioned action in a real system through a defined interface you control
   c. Run faster
   d. Cost less

**11. Choosing a primary estate after Module 5 means…**

   a. You are permanently locked out of ever touching either of the other two estates again
   b. One estate gets the deeper lab and more integration weight
   c. The other two estates are completely ignored for the rest of the course
   d. You must pick the datacenter

**12. The production-readiness checklist should be…**

   a. A document you fill in once at the very end of the engagement and then archive away
   b. A living document you update every module from M5 onward
   c. A vendor form
   d. A one-page summary

---

## Phase 4 — Build the Estate (weeks 6–8)

### Module 6 — Datacenter & Cloud: AIOps, Observability & Capacity

**Guiding question:** How do I make the datacenter's telemetry tell me things instead of just alarming — without hiding a real incident?

**Outcome:** Apply the M5 primitives to compute, storage, virtualization and cloud ops: event correlation and alert-noise reduction, anomaly detection on metrics, and capacity forecasting. Integrate with the existing monitoring rather than replacing it, and validate that "fewer alerts" did not suppress something real.

**Estate lens:** The datacenter is the best first estate because it makes telemetry, compute, storage, patching, observability and operational reliability concrete. It is also where alert fatigue is worst, so noise reduction lands immediately — provided you can prove you did not silence a genuine incident. That validation, not the correlation itself, is the FDE move.

**Apply-at-work mission — Datacenter lab: correlation, forecast & validation:** Guided lab against sample Cranfield DC telemetry: correlate an alert storm down to root events, build a simple capacity forecast for DC-1's power/cooling constraint, and produce a validation note proving no real incident was suppressed. Deliver the DC runbook artifact. (Deeper build if Datacenter is your primary estate; optional stretch: apply to your own estate.)

**Reflection:** What did correlation catch that I would have missed — and what, if anything, did it nearly hide? How confident am I the forecast is honest?

#### Resources

- [Google SRE Book — Practical alerting from time-series data](https://sre.google/sre-book/practical-alerting/) — How to turn noisy time-series into alerts that mean something — the discipline behind reducing DC-1's alert storm without going blind to real incidents.
  - Design alerts that carry signal, not just volume
  - Reason about what a quieter alert stream might hide
  - Set thresholds you can defend to the NOC
- [Elastic — What is anomaly detection?](https://www.elastic.co/what-is/anomaly-detection) — A grounding in metric anomaly detection — where it beats a static threshold and where it produces false alarms you will have to explain.
  - Explain when anomaly detection beats a fixed threshold
  - Anticipate the false positives it will generate
  - Decide which DC metrics are worth watching this way
- [Alert correlation & noise reduction (AIOps) — curated search](https://www.google.com/search?q=alert+correlation+noise+reduction+AIOps+root+cause) — Read two credible pieces on collapsing many alerts into few root events. Watch for how each claims to avoid suppressing a genuine incident.
  - Describe how correlation collapses alerts to root events
  - Interrogate a vendor's "no incident hidden" claim
  - Design a check that a real incident still surfaces
- [Data-center capacity planning (power & cooling) — curated search](https://www.google.com/search?q=data+center+capacity+planning+power+cooling+forecasting) — Background for DC-1's power/cooling constraint during the colo migration — enough to build an honest, uncertainty-aware capacity forecast.
  - Identify the constraints a DC capacity forecast must model
  - Build a forecast that states its own uncertainty
  - Avoid the single-optimistic-scenario trap
- [Datadog — What is observability?](https://www.datadoghq.com/knowledge-center/observability/) — Integrating with existing monitoring rather than replacing it — the vocabulary for making the DC pilot enrich Cranfield's stack, not compete with it.
  - Frame the pilot as integrating, not rip-and-replace
  - Map the DC signals worth feeding an AI layer
  - Preserve the history and trust the team already has

#### In-world ticket queue

> Datacenter build. DC-1 threw 900 alerts overnight — a handful real, the rest noise — and it is
> tight on power and cooling with the colo migration only half done. Dilan remembers the last "AIOps" tool that
> just moved the noise around. Prove yours reduces it without hiding a real incident.

| Ref | Priority | From | Request |
| --- | --- | --- | --- |
| CRN-0051 | P1 | Datadog → NOC | DC-1 alert storm: correlate 900 overnight alerts to root events |
| CRN-0052 | P2 | Dilan Petrou | Capacity: will DC-1 power/cooling hold through the migration? |
| CRN-0053 | P1 | Marisol Vega | Prove no real incident was suppressed by the noise reduction |

#### Project — Datacenter lab: correlation, forecast & validation (Capstone component)

Guided lab against sample Cranfield datacenter telemetry. Correlate the DC-1 overnight alert storm down to its few root events; build a simple, uncertainty-aware capacity forecast for DC-1's power/cooling constraint through the colo migration; and — the FDE move — write a validation note proving the noise reduction did not suppress a real incident. Integrate with the existing monitoring rather than replacing it. Deliver the datacenter runbook artifact. (Deeper build if Datacenter is your primary estate; optional stretch: apply the method to your own estate.)

**Deliverable:** fde/m06-datacenter-runbook.md — the correlation result, the capacity forecast with stated uncertainty, the incident-not-suppressed validation note, and the DC runbook. Updates the production-readiness checklist.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Correlation to root events | 25% | The alert storm is collapsed to a small set of root events with the reasoning shown, not just a lower count. |
| Honest capacity forecast | 25% | The forecast states its assumptions and plausible error range, and addresses DC-1's real power/cooling constraint. |
| Incident-not-suppressed validation | 35% | A concrete check demonstrates no genuine incident was hidden by the noise reduction — the safety proof, not an afterthought. |
| Integration over replacement | 15% | The pilot enriches the existing monitoring patchwork rather than proposing to rip and replace it. |

#### Scenario drills

**Drill 1.** DC-1 threw ~900 alerts overnight; a handful are real (CRN-0051), and Dilan remembers the last tool that just moved the noise around.

**Task:** Describe how you would collapse the storm to root events AND the one check you would run to prove a genuine incident was not buried. State what you would show Dilan.

**Drill 2.** DC-1 is tight on power and cooling with the colo migration half done (CRN-0052).

**Task:** Sketch a capacity forecast that Dilan could act on, naming its key assumption and the range of error you would attach to it.

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.**

**1. Why is the datacenter a good first estate to build in?**

   a. Because the datacenter is by far the cheapest estate to run any kind of experiment in
   b. It makes telemetry and reliability concrete
   c. Because the datacenter team is usually the most enthusiastic about adopting new tools
   d. Because the other two estates depend entirely on datacenter work being finished first

**2. The real FDE move in alert-noise reduction is…**

   a. Fewer alerts
   b. Proving that reducing the noise did not quietly suppress a single genuine incident
   c. A nicer dashboard
   d. Faster paging

**3. Event correlation aims to…**

   a. Permanently delete every alert that the on-call engineer finds even slightly annoying
   b. Collapse many alerts into their few root events
   c. Route each alert to a different team automatically based on some keyword match
   d. Silence P3s

**4. Anomaly detection on metrics is most valuable when…**

   a. When the operations team has completely run out of any other ideas left to try at all
   b. It catches deviations a static threshold would miss
   c. When the vendor demo looked good
   d. When alerts are already quiet

**5. A capacity forecast for DC-1 must above all be…**

   a. Presented to the CFO in a polished slide format with confident projections attached
   b. Honest about its own uncertainty
   c. Built exclusively from the single most optimistic growth scenario available to you
   d. Signed off by the vendor before it can ever be shown to anyone at all at Cranfield

**6. You integrate with existing monitoring rather than replacing it because…**

   a. It is cheaper
   b. Rip-and-replace destroys the trust and history the team already has, and it rarely lands
   c. The vendor says so
   d. It is faster

**7. The four golden signals mainly help you…**

   a. Decide precisely how many separate monitoring vendors the client really ought to buy
   b. Judge whether an alert stream carries real signal
   c. Set the exact budget for the whole observability stack for the coming year
   d. Pick a colour scheme

**8. The incident-not-suppressed validation note matters because…**

   a. Because Dilan will otherwise simply refuse to even look at the pilot's results at all
   b. "Fewer alerts" is worthless if it hid a real one
   c. It fills the runbook
   d. It pleases the NOC

**9. Correlating the alert storm down to root events means…**

   a. Manually reading through all nine hundred overnight alerts one single alert at a time
   b. Finding the few causes behind the many
   c. Escalating the entire storm straight to the most senior engineer on call that night
   d. Turning the affected monitoring system off completely until the morning shift arrives

**10. A capacity forecast is honest when…**

   a. It looks confident
   b. It states its assumptions and how wrong it could plausibly be, not just a single tidy number
   c. It is short
   d. It agrees with Dilan

**11. A real incident hidden by noise reduction is…**

   a. A perfectly acceptable and entirely normal trade-off to make in exchange for less noise
   b. The exact failure the validation guards against
   c. A problem only for the vendor to worry about later
   d. Impossible

**12. The datacenter runbook artifact should show…**

   a. A single impressive screenshot of the correlation tool working nicely during the live demo
   b. the correlation, the forecast, and the validation together
   c. Just the alert count
   d. The vendor logo

### Module 7 — Network Infrastructure: Assurance, Anomaly & Change Safety

**Guiding question:** Can AI find the "random" branch slowdown before the NOC does — and never make an unsafe change?

**Outcome:** Apply AI to network telemetry for assurance and anomaly across sites, add config-drift detection and pre-change validation, and enable natural-language query over network state — with hard guardrails keeping every change human-approved.

**Estate lens:** Network changes are the highest-blast-radius action in the estate, and Marisol's change control is usually right to say no. The FDE earns her trust by respecting the CAB, not bypassing it: AI proposes and explains a change; a human disposes. Get that boundary wrong once and you lose the room for the rest of the engagement.

**Apply-at-work mission — Network lab: root-cause & human-gated change:** Guided lab against sample Cranfield network telemetry and configs: trace the branch-slowdown to a cause using telemetry + natural-language query, detect a config drift, and propose (never auto-apply) a fix with a validation checklist and an explicit human-approval gate. Deliver the network runbook artifact. (Deeper if Network is your primary estate.)

**Reflection:** Where was I tempted to let the agent act rather than propose? What would have to be true before any network change could be safely automated here?

#### Resources

- [Atlassian — Change management (ITSM)](https://www.atlassian.com/itsm/change-management) — The change-advisory-board discipline Marisol guards. Respecting it — not bypassing it — is how AI-proposed changes earn a path to production.
  - Explain the role of a CAB in safe change
  - Position AI as proposing within change control, not around it
  - Describe how a change earns approval
- [Network assurance & intent-based networking — curated search](https://www.google.com/search?q=network+assurance+intent-based+networking+anomaly) — Read two credible pieces on AI-assisted network assurance and anomaly across sites — the pattern behind chasing Cranfield's "random" branch slowdown.
  - Describe AI-assisted assurance across many sites
  - Judge which network signals are worth correlating
  - Separate genuine assurance from vendor hype
- [Red Hat — What is configuration management?](https://www.redhat.com/en/topics/automation/what-is-configuration-management) — The grounding for config-drift detection — spotting where live config has diverged from intended before a change freeze exposes it.
  - Define configuration drift and why it accumulates
  - Design drift detection against an intended baseline
  - Treat drift as an early warning before a freeze
- [Google SRE Workbook — Canarying releases](https://sre.google/workbook/canarying-releases/) — Blast-radius thinking for change: staged, reversible rollout. Directly applicable to why a network change must never be auto-applied estate-wide.
  - Reason about blast radius for a network change
  - Design a staged, reversible change path
  - Explain why estate-wide auto-apply is unsafe
- [Natural-language query over network state — curated search](https://www.google.com/search?q=natural+language+query+network+state+troubleshooting+LLM) — How an engineer can ask the estate questions in plain English — an assistive layer that must never cross into applying changes on its own.
  - Describe NL query as an assistive, read-only layer
  - Keep query strictly separate from action
  - Spot where a plain-English answer could mislead

#### In-world ticket queue

> Network build. The "random" branch slowdown is back — three sites, no root cause after weeks — and a
> shipping-peak change freeze starts Friday. Marisol will consider an AI assist, but nothing it suggests touches
> the network without her sign-off.

| Ref | Priority | From | Request |
| --- | --- | --- | --- |
| CRN-0061 | P1 | Marisol Vega | Root-cause the "random" slowdown across 3 branches |
| CRN-0062 | P1 | Network ops | Config drift detected pre-freeze — validate before any change |
| CRN-0063 | P1 | Marisol Vega | Every AI-proposed change stays human-approved — show the gate |

#### Project — Network lab: root-cause & human-gated change (Capstone component)

Guided lab against sample Cranfield network telemetry and configs. Trace the "random" branch slowdown to a cause using telemetry plus natural-language query; detect a config drift ahead of the shipping-peak freeze; and propose — never auto-apply — a fix with a validation checklist and an explicit human-approval gate. Every AI-proposed change stays human-approved. Deliver the network runbook artifact. (Deeper build if Network is your primary estate.)

**Deliverable:** fde/m07-network-runbook.md — the root-cause trace, the drift finding, and the proposed fix with its validation checklist and human-approval gate. Updates the production-readiness checklist.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Root-cause trace | 30% | The branch slowdown is traced to a plausible cause with the telemetry/query reasoning shown, not asserted. |
| Config-drift finding | 20% | A drift from intended config is detected and framed as an early warning before the freeze. |
| Human-gated change design | 35% | The fix is proposed with evidence and an explicit approval gate; nothing is auto-applied — the boundary Marisol trusts. |
| Validation checklist | 15% | A pre-change validation checklist makes the proposed change safe to review and reversible. |

#### Scenario drills

**Drill 1.** The "random" branch slowdown is back across three sites with no root cause after weeks, and a change freeze starts Friday (CRN-0061).

**Task:** Outline how you would use telemetry + natural-language query to trace it, and state clearly where the AI stops and a human takes over.

**Drill 2.** Config drift is detected pre-freeze (CRN-0062) and must be validated before any change.

**Task:** Write the human-approval gate for the proposed fix: what the AI must show, who approves, and the hard stop if validation fails.

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.**

**1. Network changes are best understood as…**

   a. Roughly the same risk level as any other routine change made anywhere in the estate
   b. The highest-blast-radius action in the estate
   c. Best made quickly and quietly, late at night, without troubling the change board at all
   d. Something an AI can be trusted to apply on its own once it has seen enough examples

**2. You earn Marisol's trust by…**

   a. Moving fast
   b. Respecting the change board and letting AI propose while a human always disposes
   c. Buying her tool
   d. Avoiding her

**3. Config-drift detection is…**

   a. A way to automatically roll every device back to its factory-default configuration nightly
   b. Spotting where live config has diverged from intended
   c. A replacement for the entire existing change-management process the team already runs
   d. A vendor upsell

**4. Pre-change validation means…**

   a. Waiting until well after the change has been applied to see whether anything at all broke
   b. Checking a proposed change is safe before it is applied
   c. Skipping the test
   d. Asking the vendor

**5. Natural-language query over network state lets an engineer…**

   a. Completely replace their own years of hard-won networking knowledge with a chatbot
   b. Ask the estate questions in plain English
   c. Automatically apply any fix the model happens to suggest without any further review at all
   d. Bypass the change advisory board entirely whenever they happen to be in a genuine hurry

**6. AI must never auto-apply a network change because…**

   a. It is slow
   b. A wrong change can take down many sites at once, so a human must approve first every time
   c. Marisol dislikes it
   d. It costs money

**7. AI helps root-cause the "random" branch slowdown mainly because…**

   a. Because the artificial intelligence simply always knows the correct answer to every question
   b. It can correlate telemetry a human would not connect
   c. Because it can be safely left running completely unattended overnight to fix it by itself
   d. It is cheap

**8. A change freeze before the shipping peak means…**

   a. That this is exactly the perfect moment to push through as many big changes as possible
   b. Extra caution: validate, do not risk the peak
   c. That all AI work must stop
   d. Nothing much

**9. The human-approval gate on a network change is…**

   a. An optional extra step that can be safely skipped whenever the team is under time pressure
   b. The hard stop the whole design turns on
   c. A slow bureaucratic formality that mostly exists only to protect people from blame later
   d. A weekly meeting the network operations team holds to review changes already completed

**10. When AI proposes a network fix, it must also…**

   a. Apply it fast
   b. Show its evidence and reasoning so a human can judge whether the proposed change is safe
   c. Log nothing
   d. Skip the CAB

**11. Config drift detected just before a freeze is best treated as…**

   a. Something to quietly ignore until well after the busy shipping peak has fully passed by
   b. An early warning to validate before the freeze
   c. A reason to immediately cancel the entire shipping peak and impose the freeze at once
   d. A vendor problem

**12. The network runbook artifact should include…**

   a. A confident promise that the AI will soon fully automate all network changes end to end
   b. the root-cause trace, the drift finding, and the approval gate
   c. Only the final config
   d. The vendor SLA

### Module 8 — Workplace & Endpoint: Experience, Deflection & Self-Heal

**Guiding question:** How do I cut the service desk's load and lift end-user experience without leaking data or breaking devices?

**Outcome:** Apply AI to identity, endpoint and collaboration workflows: digital-employee-experience and endpoint analytics, ticket deflection with retrieval grounded in the client's runbooks, and a bounded self-heal automation — with anonymisation and PII discipline throughout, and honest measurement of what was actually deflected.

**Estate lens:** Workplace comes last because endpoint and collaboration scenarios depend on identity, access, connectivity and device posture — the things the datacenter and network modules made concrete. It is also the most data-sensitive estate: tickets are full of names and machines, so the sanitisation discipline is not optional. "Deflected" only counts if it was not merely deflected-then-reopened.

**Apply-at-work mission — Workplace lab: deflection & self-heal:** Guided lab against sample Cranfield runbooks and tickets: build a retrieval-grounded assistant over the runbook set, design one bounded self-heal flow with explicit stop conditions, add the sanitisation step, and define an honest deflection measurement. Deliver the workplace runbook artifact. (Deeper if Workplace is your primary estate.) 🎯 Completes Capstone Milestone 3.

**Reflection:** What did I have to strip before anything reached a model, and did I nearly miss a quasi-identifier? Is my deflection metric honest, or flattering?

#### Resources

- [AWS — What is Retrieval-Augmented Generation (RAG)?](https://aws.amazon.com/what-is/retrieval-augmented-generation/) — How to ground an assistant in the client's own runbooks and knowledge base so it answers from Cranfield's reality, not the model's guesswork.
  - Explain how retrieval grounds an assistant in real content
  - Describe why grounding reduces confident wrong answers
  - Design retrieval over Priya's runbook set
- [Quasi-identifiers & re-identification — curated search](https://www.google.com/search?q=quasi-identifiers+re-identification+risk+anonymisation) — Why stripping obvious names is not enough — combinations of ordinary fields can re-identify a person. The discipline for sanitising identity-dense ticket data.
  - Define a quasi-identifier and its re-identification risk
  - Spot quasi-identifiers hiding in ticket and endpoint data
  - Design a sanitisation step that catches combinations
- [Digital employee experience (DEX) — curated search](https://www.google.com/search?q=digital+employee+experience+DEX+endpoint+analytics) — The workplace-estate opportunity space: endpoint analytics and experience signals that let you lift end-user experience, not just close tickets.
  - Name the DEX and endpoint signals worth using
  - Connect experience data to a real service-desk pain
  - Separate experience improvement from ticket theatre
- [HDI / KCS — knowledge-centered service — curated search](https://www.google.com/search?q=knowledge+centered+service+KCS+ticket+deflection) — The knowledge discipline behind honest ticket deflection — and why "deflected" must mean resolved, not merely deflected-then-reopened.
  - Explain knowledge-centered service for deflection
  - Define an honest deflection metric
  - Track reopened tickets, not just first-touch closes
- [Autonomous endpoint self-heal — curated search](https://www.google.com/search?q=autonomous+endpoint+management+self-healing+guardrails) — How bounded self-heal automations work on endpoints — and why explicit stop conditions and blast-radius limits are the whole safety story.
  - Design a bounded self-heal flow with stop conditions
  - Reason about self-heal blast radius on endpoints
  - Decide when a flow must halt and hand off to a human

#### In-world ticket queue

> Workplace build. Priya's queue is drowning — 18 password resets, 12 VPN, Outlook, onboarding — and
> she is ready to pilot deflection today. Ken's only condition: no user data leaks into a model. Measure what is
> actually deflected, not what merely bounced back later.

| Ref | Priority | From | Request |
| --- | --- | --- | --- |
| CRN-0071 | P1 | Priya Raman | Ticket deflection pilot over Priya's runbooks/KB |
| CRN-0072 | P2 | you, to yourself | One bounded self-heal flow with explicit stop conditions |
| CRN-0073 | P1 | Ken Adeyemi | Ken: "prove no PII reaches the model" before go-live |

#### Project — Workplace lab: deflection & self-heal (Capstone component)

Guided lab against sample Cranfield runbooks and tickets. Build a retrieval-grounded assistant over Priya's runbook set; design one bounded self-heal flow with explicit stop conditions; add the sanitisation step that strips the identity payload (including quasi-identifiers) before anything reaches a model; and define an honest deflection measurement that counts resolution, not deflected-then-reopened. Deliver the workplace runbook artifact. (Deeper build if Workplace is your primary estate.) Completes Capstone Milestone 3.

**Deliverable:** fde/m08-workplace-runbook.md — the grounded assistant design, the bounded self-heal flow with stop conditions, the sanitisation step, and the honest deflection metric. Updates the production-readiness checklist; completes MS3.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Grounded assistant | 25% | Answers come from Cranfield's own runbooks via retrieval, with a clear reason grounding reduces made-up answers. |
| Sanitisation discipline | 30% | The identity payload — including quasi-identifiers, not just names — is stripped before any data reaches the model. |
| Bounded self-heal | 25% | One self-heal flow has explicit stop conditions and a hand-off to a human — bounded, not free-roaming. |
| Honest deflection metric | 20% | Deflection is measured as genuine resolution, explicitly tracking reopened tickets rather than first-touch bounces. |

#### Scenario drills

**Drill 1.** Priya's queue is drowning and she is ready to pilot deflection today; Ken's only condition is that no user data leaks into a model (CRN-0071, CRN-0073).

**Task:** Describe the sanitisation step you would put before the model and name one quasi-identifier you would strip that is not an obvious name.

**Drill 2.** You are designing one bounded self-heal flow with explicit stop conditions (CRN-0072).

**Task:** Define the flow's single job, its two hardest stop conditions, and how you would measure deflection honestly rather than flatteringly.

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.**

**1. The workplace is best understood as the estate that is most…**

   a. Workplace is by a wide margin the very easiest estate to gather clean data from safely
   b. data-sensitive
   c. Workplace is the estate where any mistakes carry the smallest possible consequences of all
   d. Workplace is the estate that depends least on any of the other estate work being done first

**2. Sanitisation before data reaches a model is…**

   a. Optional here
   b. Non-negotiable, because tickets are full of names, machines and other identity payload
   c. A vendor feature
   d. Ken's job only

**3. A retrieval-grounded assistant answers from…**

   a. Whatever the underlying base model happens to already know from its training data alone
   b. the client's own runbooks and knowledge base
   c. A large public dataset scraped from across the whole of the open internet somewhere
   d. Pure guesswork

**4. A ticket only counts as "deflected" if it…**

   a. If the ticket simply bounced off the assistant and then quietly came straight back later on
   b. was actually resolved, not merely deflected-then-reopened
   c. was logged
   d. was urgent

**5. A bounded self-heal flow needs above all…**

   a. A completely free hand to take whatever remediation action it judges best in the moment
   b. explicit stop conditions
   c. The single largest and most capable model that the budget could possibly stretch to cover
   d. Full unattended write access across every managed endpoint in the entire workplace estate

**6. A quasi-identifier is dangerous because…**

   a. It is loud
   b. On its own it seems harmless, but combined with other fields it can re-identify a person
   c. It is rare
   d. It is old

**7. The workplace estate is built last because…**

   a. Because it is comfortably the least important of the three estates to Cranfield overall
   b. It depends on identity, access and device posture
   c. Because the service desk team specifically asked to be scheduled into the very last slot
   d. It is hardest

**8. Grounding an assistant in real runbooks mainly…**

   a. It completely and permanently removes any and all possibility of the model ever being wrong
   b. reduces confident, made-up answers
   c. Makes it faster
   d. Costs less

**9. An honest deflection metric must also track…**

   a. Only the raw count of tickets that the assistant managed to close on its very first try
   b. reopened tickets too
   c. The total number of separate conversations the assistant had with users over the week
   d. How positive the wording of every single user feedback comment happened to sound overall

**10. The sanitisation step should strip…**

   a. Only full names
   b. the whole identity payload — hostnames, usernames, IPs and the quasi-identifiers around them
   c. The timestamps
   d. The ticket ID

**11. A self-heal flow that hits its stop condition should…**

   a. Keep quietly retrying the very same failed action over and over until it eventually works
   b. halt and hand off to a human
   c. Silently give up and close the ticket as resolved without telling anyone at all about it
   d. Escalate to the vendor

**12. The workplace runbook artifact should show…**

   a. A single glowing quote from one happy end user about how much they liked the new assistant
   b. the grounded assistant, the self-heal flow, and the sanitisation step
   c. Just the ticket count
   d. The tool logo

---

## Phase 5 — Integrate & Land (weeks 9–10)

### Module 9 — Integrated Agentic Runbooks & Production Readiness

**Guiding question:** How do I combine the estate pilots into one production-grade workflow that is monitored, reversible, and safe to run?

**Outcome:** Integrate the pilots into cross-estate agentic runbooks (reason → propose → human approves → act → verify), connected to ITSM/monitoring/config via APIs and webhooks with least-privilege accounts. Prove production readiness: eval evidence, observability dashboards, rollback, change records, and a handover pack — the M5 foundations, now demonstrated at integration.

**Estate lens:** This is synthesis, not first exposure — the runbook, tool-calling, eval, guardrail and observability foundations were taught in Module 5 precisely so this module can integrate rather than introduce. The bar here is production-readiness: an agent that touches production must propose-and-verify behind a human gate, log every action, and roll back cleanly. If it cannot, it is a demo, not a delivery.

**Apply-at-work mission — Integrated runbook pack + production-readiness review:** Wire one pilot to ServiceNow (ticket in → AI triage → drafted action → human approve → verify) and build one cross-estate agentic runbook with an explicit stop-and-approve. Run the production-readiness checklist against it — test evidence, rollback, observability, HITL, change record, handover — and deliver the integrated runbook pack with observability evidence and the HITL gate design. 🎯 Completes Capstone Milestone 4.

**Reflection:** Which production-readiness item was my runbook missing when I first checked? Where is the single point where a wrong AI step could still cause harm?

#### Resources

- [Martin Fowler — Engineering practices for LLM applications](https://martinfowler.com/articles/engineering-practices-for-llm.html) — What it takes to run an LLM feature as a real system — testing, evals, observability, guardrails — the production bar this module holds the integrated runbook to.
  - Name the engineering practices a production AI feature needs
  - Judge an integrated runbook against a production bar
  - Spot where a pilot is still only a demo
- [ServiceNow REST API integration & webhooks — curated search](https://www.google.com/search?q=ServiceNow+REST+API+integration+webhook+least+privilege) — How to wire a workflow to ServiceNow with scoped, least-privilege access — the plumbing for "ticket in → AI triage → human approve → verify".
  - Describe a least-privilege ServiceNow integration
  - Design the ticket → triage → approve → verify loop
  - Keep write access scoped to exactly the workflow
- [Safe rollback & deployment strategies — curated search](https://www.google.com/search?q=safe+rollback+deployment+strategies+reversible+changes) — The discipline of reversible change — so that when an AI-driven action goes wrong, you can undo it cleanly. A production-readiness must-have.
  - Design a clean rollback for an AI-driven action
  - Reason about which actions are reversible and which are not
  - Make reversibility a precondition for autonomy
- [Human-in-the-loop AI in production — curated search](https://www.google.com/search?q=human+in+the+loop+AI+production+approval+workflow) — Patterns for keeping a human approval point in a live workflow without grinding it to a halt — the balance the integrated runbook has to strike.
  - Place a human approval point without killing throughput
  - Decide which actions need a gate and which do not
  - Design the approver's view so the decision is real
- [Production readiness review — curated search](https://www.google.com/search?q=production+readiness+review+checklist+SRE) — The SRE practice of a production-readiness review — the model for running your living checklist against the runbook before it goes live.
  - Run a production-readiness review against a runbook
  - Turn the living checklist into a go/no-go
  - Find the readiness item the runbook is missing

#### In-world ticket queue

> Integration week. The three pilots work in isolation; now they must become one production-grade
> workflow wired to ServiceNow, monitored, and reversible. Ken will not approve anything that acts on production
> without a human gate, an audit trail, and a clean rollback. This is where "demo" becomes "delivery".

| Ref | Priority | From | Request |
| --- | --- | --- | --- |
| CRN-0081 | P1 | you, to yourself | Wire a pilot to ServiceNow: ticket → AI triage → human approve → verify |
| CRN-0082 | P1 | Renata Cole | Run the production-readiness checklist against the runbook |
| CRN-0083 | P1 | Ken Adeyemi | Ken: "show me rollback, audit trail and the human gate" |
| CRN-0084 | P2 | Marisol Vega | Observability: dashboards + alerts for the AI workflow itself |

#### Project — Integrated runbook pack + production-readiness review (Capstone component)

Wire one pilot to ServiceNow (ticket in → AI triage → drafted action → human approve → verify) with least-privilege access, and build one cross-estate agentic runbook with an explicit stop-and-approve step. Run the production-readiness checklist against it — eval evidence, rollback, observability, HITL gate, change record, handover — and deliver the integrated runbook pack with observability evidence and the HITL gate design. This is synthesis of the Module-5 foundations at integration, not first exposure. Completes Capstone Milestone 4.

**Deliverable:** fde/m09-integrated-runbook.md — the wired ServiceNow loop, the cross-estate runbook with its stop-and-approve, the completed production-readiness review, and the observability + HITL evidence. Completes MS4.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Integrated, gated workflow | 30% | One pilot is wired end to end (ticket → triage → approve → verify) with an explicit human gate before any action. |
| Production-readiness evidence | 30% | Eval evidence, rollback, observability and change record are demonstrated against the living checklist, not merely claimed. |
| Least-privilege integration | 20% | API/webhook access is scoped to the workflow with least privilege — the integration Ken would sign. |
| Synthesis, not re-teaching | 20% | The runbook applies the Module-5 primitives at integration and names the single point where a wrong step could still cause harm. |

#### Scenario drills

**Drill 1.** You must wire a pilot to ServiceNow: ticket in → AI triage → human approve → verify (CRN-0081), and Ken wants rollback, audit trail and the human gate shown (CRN-0083).

**Task:** Describe the least-privilege access you would request and where exactly the human gate sits. Name the one action in the loop that must never happen without approval.

**Drill 2.** You are running the production-readiness checklist against the runbook (CRN-0082).

**Task:** List the readiness items you would check and identify the one your runbook is most likely to be missing on the first pass.

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.**

**1. Module 9 is best understood as…**

   a. An entirely fresh introduction to tool calling, evals, guardrails and observability from scratch
   b. synthesis, not first exposure
   c. The first time the learner has ever had to think about production readiness at all here
   d. A lighter, optional module that mostly just reviews what the earlier modules already covered

**2. An agentic runbook near production follows the pattern…**

   a. Act, then log
   b. Reason, propose, a human approves, then act, and finally verify the result afterwards
   c. Just act
   d. Ask, then wait

**3. "Production readiness" here means…**

   a. That the demo ran successfully end to end at least once in front of the assembled steering group
   b. eval evidence, observability, rollback, and a human gate
   c. That the vendor has formally certified the whole thing as being ready for production use somewhere
   d. That it is fast

**4. The integration to ServiceNow should use…**

   a. A single shared administrator account with full write access to absolutely everything at once
   b. least-privilege API access scoped to the workflow
   c. The vendor default
   d. A personal login

**5. The real difference between a demo and a delivery is that…**

   a. A delivery generally has a noticeably larger and more detailed accompanying slide deck than a demo
   b. a demo cannot be safely run in production
   c. A delivery is the one that the client happens to have paid a slightly larger sum of money for
   d. A delivery is simply a demo that has been given a second time to a different audience later

**6. A clean rollback matters because…**

   a. It looks tidy
   b. When an AI-driven action goes wrong, being able to reverse it cleanly is what limits the harm
   c. The vendor requires it
   d. It is quick

**7. Observability dashboards for the AI workflow itself let the ops team…**

   a. Finally replace every one of the older monitoring dashboards the team currently still relies on
   b. see what the AI is doing and catch drift
   c. Prove to the finance team that the whole project was genuinely worth all the money spent on it
   d. Look busy

**8. A change record for an AI action exists so that…**

   a. So that there is always someone specific who can be held personally to blame when it goes wrong
   b. every AI-driven change is auditable after the fact
   c. It fills the log
   d. The CAB is happy

**9. The handover pack matters because…**

   a. Because the delivery lead insists on having one final document produced before the engagement ends
   b. the client's team must run it after you leave
   c. Because it is a standard contractual deliverable that every engagement of this kind must include
   d. Because it gives the FDE something tangible to point at when justifying the final invoice later

**10. The single most dangerous point in an integrated runbook is…**

   a. The slowest step
   b. any point where a wrong AI step could still act on production without a human catching it first
   c. The first step
   d. The log step

**11. At integration, the Module-5 foundations are…**

   a. Being taught to the learner for the very first time, now that the estate labs have all finished
   b. demonstrated in production, not introduced
   c. Quietly abandoned in favour of whatever the integration vendor happens to recommend doing instead
   d. Optional now

**12. A cross-estate agentic runbook needs an explicit stop-and-approve because…**

   a. Because the change advisory board will otherwise be quite annoyed about being left out of it all
   b. no agent should act across estates unattended
   c. It is neater
   d. The demo needs it

### Module 10 — Governance, Adoption, Scale & Capstone Defense

**Guiding question:** How do I make this outlast me — governed, adopted by the client's team, measured, and defended to a sceptical room?

**Outcome:** Consolidate the governance thread into a signable AI-in-operations pack (data residency, change control for AI actions, guardrails, audit, the autonomy ladder). Drive adoption and handover to the client's ops team, measure outcomes and ROI against the Module-2 baseline, tell the demo narrative, plan the land-and-expand, and defend the whole engagement to the Cranfield cast.

**Estate lens:** An engagement that dies when the FDE leaves was theatre. Landing means the client's own team runs it, the CISO and change board have signed the governance, the CFO has a defensible number, and there is a plan to expand from pilot to estate. The capstone is not a presentation — it is the defence of accumulated work in front of people paid to be sceptical.

**Apply-at-work mission — Capstone: defend the Cranfield engagement:** Assemble the accumulated engagement pack — charter, estate map + baseline, use-case canvas, blueprint + readiness gate, delivery foundation, the estate runbooks, the integrated runbook pack, and the governance/adoption pack — and defend it to the cast with measured before/after. Choose a track: (A) delivery-lead package (the full engagement, presentation-ready) or (B) a deep single-estate implementation with a working integration.

**Reflection:** Reopen the Module-1 sealed baseline: what did I get wrong on day one? What would I need to be true before I let any of this run unattended?

#### Resources

- [ISO/IEC 42001 & enterprise AI governance — curated search](https://www.google.com/search?q=ISO+IEC+42001+AI+management+system+governance) — The shape of a formal AI management system — the reference for turning your threaded governance work into a signable AI-in-operations pack.
  - Describe what a formal AI governance pack contains
  - Map your threaded controls to a signable structure
  - Speak the governance language a CISO expects
- [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework) — The autonomy-and-controls backbone for the governance pack: risk tiers, controls, and the change control that governs AI actions in production.
  - Frame the autonomy ladder in risk terms
  - Attach controls to each rung of autonomy
  - Govern AI-driven change like any other change
- [Prosci — ADKAR change-management model](https://www.prosci.com/methodology/adkar) — Adoption is a people problem, not a technical one. ADKAR is a clean frame for handover and getting the client's own team to actually run it.
  - Plan adoption as a change-management problem
  - Design a handover the client's team will sustain
  - Anticipate where adoption usually stalls
- [Measuring ROI of IT-operations automation — curated search](https://www.google.com/search?q=measuring+ROI+IT+operations+automation+baseline) — How to build a defensible ROI number against a real baseline — what Tom will fund, as opposed to a slideware projection.
  - Build an ROI number against the Module-2 baseline
  - Separate a defensible figure from a flattering one
  - Present ROI a sceptical CFO will actually accept
- [Land-and-expand strategy — curated search](https://www.google.com/search?q=land+and+expand+strategy+pilot+to+production+scale) — How one proven pilot becomes a path across the estate — the 90-day plan Renata wants to turn this engagement into a repeatable play.
  - Turn a proven pilot into a staged expansion plan
  - Sequence the estate rollout by value and readiness
  - Frame expansion around evidence, not enthusiasm

#### In-world ticket queue

> Final week. Steering review Friday: Dilan wants to know it will outlast you, Tom wants the number
> against the Week-2 baseline, Ken wants a signable governance pack, and Renata wants a handover the client's own
> team can run. Defend the whole engagement — and make the case to expand from pilot to estate.

| Ref | Priority | From | Request |
| --- | --- | --- | --- |
| CRN-0091 | P1 | Dilan Petrou | Capstone defense: architecture, evidence, governance, adoption plan |
| CRN-0092 | P1 | Tom Böttcher | Tom: "the number, measured against where we started" |
| CRN-0093 | P1 | Ken Adeyemi | Signable AI-in-operations governance pack |
| CRN-0094 | P2 | Renata Cole | Handover pack + 90-day land-and-expand plan |

#### Project — Capstone: defend the Cranfield engagement (Capstone component)

Assemble the accumulated engagement pack — charter, estate map + baseline, opportunity map, blueprint + readiness gate, delivery foundation, the three estate runbooks, the integrated runbook pack, and the governance/adoption pack — and defend it to the Cranfield cast with a measured before/after against the Module-2 baseline. Consolidate governance into a signable AI-in-operations pack (data residency, change control for AI actions, guardrails, audit, the autonomy ladder), plan the handover and the 90-day land-and-expand, and reopen the Module-1 sealed baseline. Choose a track: (A) delivery-lead package — the full engagement, presentation-ready; or (B) a deep single-estate implementation with a working integration.

**Deliverable:** fde/m10-capstone.md — the assembled engagement pack, the signable governance pack, the measured ROI against the Module-2 baseline, the adoption + land-and-expand plan, and the reopened sealed baseline. The defended whole.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Measured before/after | 25% | Outcomes are shown against the dated Module-2 baseline as a defensible number, not a feeling or a projection. |
| Signable governance pack | 25% | Data residency, AI change control, guardrails, audit and the autonomy ladder are consolidated into something the CISO and CAB could actually sign. |
| Adoption & land-and-expand | 25% | A handover plan that lets the client's team run it, plus a 90-day expansion path sequenced by value and readiness. |
| Honest reflection | 25% | The Module-1 sealed baseline is reopened and what changed is named honestly — the growth, not a polished restatement. |

#### Scenario drills

**Drill 1.** Steering review Friday: Dilan wants proof it will outlast you, Tom wants the number against the baseline, Ken wants a signable pack (CRN-0091, CRN-0092, CRN-0093).

**Task:** For each of the three, state the one artifact you would put in front of them and the single sentence you would open with.

**Drill 2.** Renata wants a handover pack and 90-day land-and-expand plan the client's own team can run (CRN-0094).

**Task:** Outline the first three expansion steps after the pilot and the one condition that must be true before each, so expansion is evidence-led not hope-led.

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/fde-infrastructure to check yourself.**

**1. An engagement that dies the moment the FDE leaves was…**

   a. A perfectly acceptable and entirely normal outcome for this kind of short embedded engagement
   b. theatre
   c. Actually a genuine success, because the FDE clearly made themselves completely indispensable to it
   d. Mostly the client's own fault for not hiring enough of their own permanent staff to run it later

**2. "Landing" the engagement means…**

   a. A nice final demo
   b. The client's own team runs it, governance is signed, and there is a plan to expand it further
   c. A thick report
   d. A closing dinner

**3. The governance pack must be…**

   a. A lengthy academic document that carefully surveys the entire field of AI ethics in general terms
   b. signable by the CISO and the change board
   c. Written entirely by the external vendor and then simply handed over to the client to sign off on
   d. Kept confidential

**4. ROI must be measured…**

   a. However the chief financial officer happens to prefer to see it presented on the day in question
   b. against the dated Module-2 baseline
   c. By the vendor
   d. Optimistically

**5. The autonomy ladder is best understood as something…**

   a. A one-off decision the client makes right at the very start about exactly how autonomous AI may be
   b. earned rung by rung with evidence
   c. A marketing term the vendor uses to describe how advanced their particular product apparently is
   d. A fixed regulatory requirement that applies identically to every AI system regardless of its risk

**6. Adoption by the client's team requires…**

   a. A launch email
   b. Handover, training and enough trust that the client's own people will actually keep running it
   c. A big demo
   d. A logo

**7. The capstone is best understood as…**

   a. A polished one-way presentation delivered to a friendly and broadly supportive internal audience
   b. a defence of accumulated work to sceptics
   c. A written exam that tests whether the learner has memorised the key terms from the whole course
   d. A formality

**8. Change control for AI actions means…**

   a. That every single AI action must first be reviewed in person by the full change board in advance
   b. AI-driven changes follow the same governed path
   c. It moves faster
   d. The vendor decides

**9. Reopening the Module-1 sealed baseline shows…**

   a. Whether the tutor happened to mark the very first week-one assignment fairly and consistently or not
   b. how the learner's thinking actually changed
   c. Whether the original unassisted answer was already polished enough to reuse without any real edits
   d. That the course has now finally reached its proper administrative end and can be formally closed

**10. A defensible ROI number is one that…**

   a. Sounds impressive
   b. ties a measured after-figure back to the dated before-figure, so the CFO can actually trust it
   c. Is very large
   d. Is rounded up

**11. The land-and-expand plan is valuable because…**

   a. Because the FDE's own delivery lead is keen to turn this one engagement into repeatable practice revenue
   b. it turns one proven pilot into a path across the estate
   c. Because the client will always automatically want to buy considerably more the moment the pilot works
   d. It fills a slide

**12. Data residency belongs in the governance pack because…**

   a. Because Ken personally happens to feel rather strongly about it for his own particular reasons
   b. it is one of the CISO's hard, non-negotiable lines
   c. It sounds rigorous
   d. The vendor asked

## Toolkits

### Pre-work diagnostic → remediation map

**Unlocks in module 1.**

A quick self-assessment across the six competencies, mapped to targeted "do these before Module 5" actions — not just a score.

```markdown
# Pre-work diagnostic → remediation map

Rate yourself 1–5 on each axis, then follow the remediation for anything under 3.
The point is targeted preparation before the delivery foundation (Module 5), not a grade.

| Competency | 1–5 | If under 3, do this before M5 |
|---|---|---|
| Infrastructure depth (DC / network / workplace) | | Sketch your own estate's topology and systems of record from memory, then check it |
| AI fluency in operations | | Watch the Karpathy intro (M1) and write one paragraph on what an LLM cannot do |
| Discovery & scoping | | Practise turning one real pain into a value × data-readiness × risk line |
| Hands-on delivery | | Read the Anthropic tool-use overview (M5) and the evals piece; note what you don't follow |
| Governance, risk & security | | Read the OWASP LLM Top 10; list the three risks most relevant to your estate |
| Client trust & communication | | Write the two-sentence reply to a sceptical sponsor (M1 drill) |

## How to use it
- Retake it after Module 5 and again at Module 10 — the delta is evidence for your journal.
- Bring the lowest axis into your primary-estate choice: sometimes the right estate is the
 one that stretches your weakest competency, sometimes the one that plays to your strongest.
```

### Opportunity-scoring sheet

**Unlocks in module 3.**

The value × data-readiness × risk grid, with a quantified KPI hypothesis and a buy / integrate / build call per shortlisted item.

```markdown
# Opportunity-scoring sheet

Score each candidate 1–5 on all three axes. Shortlist the highest combined, then commit a
number. Readiness comes from the Module-2 data inventory, not optimism.

| Opportunity | Value (1–5) | Data-readiness (1–5) | Risk (1–5, lower=safer) | Shortlist? |
|---|---|---|---|---|
| DC alert-noise reduction | | | | |
| Capacity forecasting | | | | |
| Network anomaly / assurance | | | | |
| Config-drift detection | | | | |
| Ticket deflection | | | | |
| Bounded self-heal | | | | |
| Security operations assist | | | | |
| DC energy optimisation | | | | |

## For each shortlisted item
- **KPI hypothesis (quantified, vs the M2 baseline):** e.g. "cut DC-1 alert noise 40% without
 suppressing a real incident."
- **Risk tier:** low / medium / high — and the one control that earns the tier.
- **Buy / integrate / build:** with a one-line rationale. Integrate is the default; build is
 the exception, reserved for specific, valuable needs nothing else fits.

## The killed idea
Name one tempting solution-first idea you rejected, and whether value or data killed it.
Recording the rejection is the evidence of triage.
```

### Solution-blueprint template

**Unlocks in module 4.**

Architecture, scope, dependencies, delivery plan and acceptance criteria — the bridge from a chosen opportunity to a buildable plan.

```markdown
# Solution-blueprint template

## 1. Problem & outcome
- The bounded problem this pilot solves (one sentence).
- The acceptance criteria: what "done and good" looks like, as checkable statements.

## 2. Architecture
- Data sources in → processing → action out. Where the AI sits and what it may touch.
- The advisory-before-action rung this pilot occupies (advise / propose / act-with-approval).

## 3. Scope
- In scope (this pilot). Out of scope (named later phases).

## 4. Dependencies
- Access grants, data feeds, sign-offs, environments. The unnamed dependency is the one
 that stalls the build in week six — write them all down.

## 5. Delivery plan
- Milestones, sequence, who does what, and the human-approval path for any action.

## 6. Blast radius
- What a wrong action could touch, and how far it could spread. The control follows the reach.
```

### Delivery-foundation checklist (living)

**Unlocks in module 5.**

The reusable data / access / eval / guardrail scaffold every estate lab depends on. Stood up once in M5, referenced everywhere after.

```markdown
# Delivery-foundation checklist (living — start in M5)

## Block A — data & access
- [ ] Data sources identified with a contract (what, shape, who may use it)
- [ ] ITSM / CMDB / monitoring API access scoped and requested
- [ ] Least-privilege service accounts specified (read-only where possible)
- [ ] Identity / RBAC model written down and signed by the CISO
- [ ] Ticket / event ingestion path defined
- [ ] Runbook anatomy understood for the target workflow

## Block B — safety & control
- [ ] Tool / function calling interface defined and bounded
- [ ] Evals written as acceptance criteria (before build)
- [ ] Guardrails specified (what the AI may and may not do)
- [ ] Observability designed (traces, metrics, logs for the AI itself)
- [ ] Human-in-the-loop gate placed before any production action
- [ ] Fallback path defined (what happens when the AI is unsure or fails)
- [ ] Failure modes reviewed against the failure-modes library

## Primary-estate choice
- [ ] Primary estate nominated (Datacenter / Network / Workplace) with a rationale
```

### Production-readiness checklist (living)

**Unlocks in module 9.**

The go/no-go for letting a runbook near production. Begun in M5, updated every module, run in full at integration (M9).

```markdown
# Production-readiness checklist (living — run in full at M9)

A runbook that cannot tick these is a demo, not a delivery.

## Evidence
- [ ] Eval evidence: the AI meets its acceptance criteria, shown not claimed
- [ ] Observability: dashboards + alerts for the AI workflow itself
- [ ] Change record: every AI-driven action is auditable after the fact

## Safety
- [ ] Human-in-the-loop gate before any production action
- [ ] Clean rollback for every action the AI can take
- [ ] Least-privilege access, scoped to the workflow
- [ ] The single point where a wrong step could still cause harm is identified and gated

## Handover
- [ ] The client's own team can operate it (runbook + training)
- [ ] Handover pack complete
- [ ] Escalation and stop conditions documented
```

### Failure-modes & anti-patterns library

**Unlocks in module 5.**

The catalogue of ways AI-in-operations goes wrong — from AI theatre to suppressed incidents to prompt injection — each with its guardrail.

```markdown
# Failure-modes & anti-patterns library

Referenced in M5 (foundation), M9 (integration), and the governance pack.

## Delivery anti-patterns
- **AI theatre** — a slick demo of an unowned, unneeded capability. *Guard:* every pilot has a
 named owner and a measured problem behind it.
- **Solution-first** — building the clever thing, then hunting for a problem. *Guard:* triage on
 value × data-readiness × risk before any build.
- **Baseline-free impact** — "it feels faster." *Guard:* a dated Module-2 baseline behind every KPI.
- **The hero engagement** — only the FDE can run it. *Guard:* handover pack + the client's team operates it.

## Operational failure modes
- **Suppressed incident** — noise reduction hides a real event. *Guard:* the not-suppressed validation note.
- **Confident wrong answer** — fluent hallucination. *Guard:* retrieval grounding + evals + a human gate.
- **Unattended production action** — AI acts without approval. *Guard:* HITL gate + advisory-before-action ladder.
- **High-blast-radius change** — one network change takes down many sites. *Guard:* propose-not-apply + CAB.

## Data & security failure modes
- **Identity leak** — names / hosts / IPs reach a model. *Guard:* the sanitisation step (incl. quasi-identifiers).
- **Prompt injection / excessive agency** (OWASP LLM Top 10). *Guard:* input handling + bounded tools + least privilege.
- **Data residency breach** — telemetry leaves its permitted region. *Guard:* mapped data locations + governance pack.
```

### Governance-pack skeleton

**Unlocks in module 10.**

The signable AI-in-operations pack the CISO and change board approve — data residency, AI change control, guardrails, audit, the autonomy ladder.

```markdown
# AI-in-operations governance pack (skeleton)

Threaded from Module 1, consolidated here into something the CISO and CAB can sign.

## 1. Scope & data residency
- Which systems and data the AI touches, and where every byte lives.

## 2. Change control for AI actions
- How an AI-proposed change is proposed, approved and recorded — the same governed path as any change.

## 3. Guardrails & autonomy ladder
- What the AI may and may not do at each rung; how a rung is earned with evidence.

## 4. Access & least privilege
- Service accounts, scopes, and who holds them.

## 5. Audit & observability
- What is logged, retained, and how an action is reconstructed after the fact.

## 6. Human-in-the-loop
- Where the gates are, who approves, and the stop conditions.

## 7. Sign-off
- CISO / DPO, change board, sponsor — names and date.
```

### Per-persona feedback templates

**Unlocks in module 1.**

Structured checkpoint prompts for each of the Cranfield cast, so client feedback is specific and comparable across modules.

```markdown
# Per-persona feedback templates

Use at each client checkpoint. Match the evidence to what the persona actually fears.

## Dilan Petrou (sponsor, burned by shelfware)
- "Show me this will outlast the engagement." → handover plan + adoption.
- Opening line: the one measurable Week-10 outcome you put your name against.

## Marisol Vega (NOC / change control)
- "Walk me through how an AI-proposed change gets approved." → the human-approval gate.
- She must see: propose-not-apply, evidence shown, an explicit stop.

## Ken Adeyemi (CISO / DPO)
- "Show me the data-access, eval and risk controls in writing." → checkable readiness-gate statements.
- He must see: data residency, least privilege, no unattended production action.

## Priya Raman (service desk, early adopter)
- "Does this genuinely help my team without leaking user data?" → sanitisation + honest deflection metric.

## Tom Böttcher (CFO / procurement)
- "What does this actually save, against where we started?" → measured before/after vs the M2 baseline.

## Renata Cole (your delivery lead)
- "Is this a repeatable land-and-expand play?" → the artifacts + the 90-day expansion plan.
```

## Capstone

### Track A — Delivery lead · Full engagement package

The complete, presentation-ready Cranfield engagement: charter, estate map + baseline, opportunity map, blueprint + readiness gate, delivery foundation, three estate runbooks, integrated runbook pack, and the governance/adoption pack — defended with measured before/after.

### Track A — Delivery lead · AI-in-operations governance standard

The standing control the client keeps: a signable governance pack — data residency, change control for AI actions, guardrails, audit, and the autonomy ladder — written so the CISO and change board actually approve it.

### Track A — Delivery lead · Land-and-expand playbook

Renata's repeatable play: how one proven Cranfield pilot becomes a path across the estate, sequenced by value and readiness, with the 90-day plan and the artifacts that make it repeatable on the next account.

### Track B — Deep build · Datacenter runbook, integrated

Go deep on one estate: the DC alert-correlation + capacity + not-suppressed validation runbook, wired to ServiceNow with a human gate, observability and rollback — a working integration, production-ready.

### Track B — Deep build · Network assurance runbook, integrated

The network root-cause + config-drift + human-gated-change runbook as a working, reversible integration — every AI-proposed change human-approved, with the evidence Marisol would need to trust it.

### Track B — Deep build · Workplace deflection runbook, integrated

The retrieval-grounded deflection + bounded self-heal runbook with the sanitisation step and an honest deflection metric, integrated end to end — proving no user data reaches the model.

### Your own · Your own estate

The strongest capstone is a real engagement you actually face. If your organisation or a client has a live AI-in-operations decision, run the method on that one — the pushback will be free and genuine.

### Milestones

- **Module 2 — Estate assessment & baseline**
- **Module 4 — Blueprint + delivery-readiness gate**
- **Module 8 — Three estate runbooks**
- **Module 9 — Integrated runbook pack**
- **Module 10 — Engagement defended**

### Portfolio checklist

- Engagement charter + the dated Module-1 sealed baseline
- Estate map + rated data inventory + measured baseline (question marks preserved then closed)
- Opportunity map with quantified KPI hypotheses and buy/integrate/build calls
- Solution blueprint + the readiness gate as checkable statements
- Delivery-foundation plan + the shared scaffold, and your primary-estate choice
- Datacenter runbook: correlation, capacity forecast, incident-not-suppressed validation
- Network runbook: root-cause trace, drift finding, human-gated change
- Workplace runbook: grounded assistant, bounded self-heal, sanitisation, honest deflection metric
- Integrated runbook pack with observability evidence and the HITL gate design
- Signable AI-in-operations governance pack
- Measured ROI against the Module-2 baseline + the 90-day land-and-expand plan
- Capstone (Track A package or Track B deep build), presented and defended
