# Agentic AI for Leadership

> Executive Certificate — 18 weeks, self-paced, modelled on the NTU Singapore Executive Certificate in Agentic AI Leadership

- **Audience:** Executives · Managers · Consultants · Product & Ops leaders
- **Level:** Beginner → Advanced (no coding required)
- **Duration:** 18 weeks · 6–10 h/week
- **Modules:** 18
- **Pass mark:** 70%
- **Interactive version:** https://ragentic.netlify.app/#/courses/agentic-ai-leadership

**This file is generated from the course data by `scripts/build-notes.mjs`. Edit the course data, not this file.**

---

## Phase 1 — Foundations: From Generative to Agentic (weeks 1–3)

### Module 1 — From Search to Digital Teammates

**Outcome:** Experience the shift from AI-as-search-tool to AI-as-collaborative-teammate — and feel how your own workflow changes.

**Leadership lens:** Your calendar is full of work a persistent teammate could carry: briefing packs, meeting prep, first drafts. The leaders who win aren't the best prompters — they're the best delegators to a new kind of colleague.

**Apply-at-work mission — Hire your first AI teammate:** Stand up a persistent AI teammate (Claude Project / Custom GPT / NotebookLM) for ONE recurring task in your actual job — give it a role, goals, and your working style. Use it every working day this week.

**Reflection:** How does my thinking change when I treat AI as a teammate instead of a tool? What did I delegate this week that I would never have delegated a month ago?

#### Resources

- [Andrej Karpathy — Intro to Large Language Models](https://www.youtube.com/watch?v=zjkBMFhNj_g) — The one-hour mental model of what these systems really are — the foundation every leadership conversation builds on.
  - Ground every leadership conversation in what an LLM actually is, not the marketing
  - Explain "AI as teammate vs AI as search" to a board in plain terms
  - Separate durable capability from hype when a vendor pitches
- [Anthropic — Research & how modern assistants are built](https://www.anthropic.com/research) — How frontier labs think about capability, safety, and where assistants are heading.
  - Speak to where frontier assistants are heading, with capability and safety in view
  - Borrow a lab's framing for a risk conversation with peers
  - Judge a roadmap claim against how builders actually talk
- [NotebookLM — official guide & examples](https://notebooklm.google/) — Your first "knowledge teammate": upload documents, then converse with a system that knows your world.
  - Stand up your first "knowledge teammate" from your own documents
  - Experience grounded assistance versus open generation directly
  - See what an AI that knows your world can and cannot do
- [Claude Projects / Custom GPTs — persistent context guides](https://support.claude.com/en/articles/9517075-what-are-projects) — The mechanics of giving an AI a standing role, goals, and memory of how you work.
  - Give an AI a standing role, goals and memory of how you work
  - Explain persistent context as the line between tool and teammate
  - Design a reusable assistant instead of one-off prompts
- [The rise of AI agents — visual explainers (curated search)](https://www.youtube.com/results?search_query=the+rise+of+ai+agents+agentic+ai) — Pick one strong explainer on the search→copilot→agent evolution; watch critically, note the hype.
  - Trace the search → copilot → agent evolution and mark the hype
  - Pick a credible explainer and separate shipped from aspirational
  - Form your own line on where agents actually are today

#### Project — Your First AI Teammate

Create a persistent AI teammate using Claude Projects or a Custom GPT: give it a clear role, goals, and knowledge of your working style. Separately, load 10–15 documents about your industry or function into NotebookLM and hold three deep conversations with your "knowledge teammate". Then write a 500-word reflection: "How my thinking changes when I treat AI as a teammate instead of a tool."

**Deliverable:** portfolio/w01-ai-teammate.md — teammate setup (role prompt), 3 conversation takeaways, and the 500-word reflection.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Teammate design | 25% | The role prompt defines role, goals, tone, and working context specifically enough that a colleague could tell whose teammate it is; not a generic "helpful assistant". |
| Real daily use | 25% | Evidence of use on 5+ real work items across the week (drafts, prep, analysis), with notes on what worked and what disappointed. |
| Knowledge-teammate experiment | 20% | 10–15 genuinely relevant documents in NotebookLM; three conversations that surface at least one insight you did not already have. |
| Reflection quality | 30% | The 500 words name a concrete change in how you think or delegate — not "AI is amazing" but a specific before/after in your own behaviour. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. What most fundamentally distinguishes an AI "teammate" from an AI "tool"?**

   a. Bigger model size
   b. Persistence, memory, context, and initiative
   c. Considerably faster average response times overall
   d. Access to the internet

**2. The evolution this module traces is best described as…**

   a. Search engines → databases → spreadsheets
   b. Search → copilots → autonomous agents
   c. Chatbots → progressively bigger chatbots
   d. On-prem → to cloud → to the edge over time

**3. Why does giving an AI a standing role and goals ("persistent context") matter for leaders?**

   a. It reduces API costs
   b. It turns each interaction into true delegation
   c. It makes the model considerably more accurate at maths
   d. It is required by vendors

**4. A leader uploads 15 strategy documents to NotebookLM and interrogates them. This practice primarily builds…**

   a. A full production-grade RAG system
   b. A personal knowledge teammate
   c. A properly fine-tuned custom model
   d. An automation pipeline of some kind

**5. Which behaviour signals someone still treats AI as a tool, not a teammate?**

   a. Reviewing all AI output before using it
   b. Giving it a clear role and feedback over time
   c. Re-explaining context each time
   d. Delegating your meeting prep to it

**6. The "experience layer" shift (voice, multimodal, always-available AI) matters to leaders mainly because…**

   a. It looks modern in demos
   b. It changes when and where thinking work happens, not just how fast
   c. It reduces licence costs
   d. It removes the need for meetings

**7. What is the biggest realistic risk of week 1's "AI teammate" practice?**

   a. The AI simply refuses to do any of the work
   b. Accepting fluent output without judgment
   c. Quietly running out of all your available tokens
   d. Colleagues noticing

**8. Why do many professionals report AI "doesn't help much" after trying it once?**

   a. The models are weak
   b. They issued one-shot queries with no context
   c. AI only helps engineers
   d. Their privacy settings quietly block the quality

**9. A well-designed teammate role prompt should include…**

   a. Only the single small task of the day itself
   b. Role, goals, audience, tone, and style
   c. As little as you possibly can, to avoid bias
   d. Technical model parameters

**10. The reflection exercise ("how my thinking changes") exists because…**

   a. Writing fills the week
   b. Leadership development happens through reflection on practice, not consumption of content
   c. It produces content for LinkedIn
   d. It tests writing skills

**11. Which task is the WEAKEST first delegation to a new AI teammate?**

   a. Drafting up a meeting agenda from your notes
   b. Summarising a rather long and detailed report
   c. A final, unreviewed client message
   d. Preparing interview questions

**12. The healthiest mental model for current AI teammates is…**

   a. An infallible oracle
   b. A fast, tireless junior colleague with confidence issues in reverse — always confident, sometimes wrong
   c. A search engine with better grammar
   d. A threat to be minimised

### Module 2 — Why Now: The AI Inflection Point

**Outcome:** Explain the technical and economic forces behind the current inflection — and why agentic AI is not another RPA wave.

**Leadership lens:** Boards ask "why now, why us, why this much?" This week gives you the cost-collapse, capability-explosion, and time-to-X arguments to answer in their language: valuation, revenue per employee, competitive moats.

**Apply-at-work mission — Brief your team on the inflection:** Run a 15-minute briefing for your team or manager: one chart, one industry example, one implication for your business in the next 12 months. Note what convinced them and what met resistance.

**Reflection:** Which part of my industry's value chain is most exposed to the time-to-X collapse — and am I positioned as a spectator or an architect of that change?

#### Resources

- [McKinsey — The economic potential of generative AI](https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier) — The report boards quote: where the trillions in value supposedly sit, function by function.
  - Cite where the value supposedly sits, function by function
  - Read the report boards quote — critically
  - Separate McKinsey's evidence from its enthusiasm
- [Stanford HAI — AI Index Report](https://aiindex.stanford.edu/) — The definitive annual dataset: capability curves, cost curves, adoption numbers. Skim the executive summary hard.
  - Pull capability, cost and adoption numbers from the definitive dataset
  - Ground an inflection-point claim in real curves
  - Quote a defensible statistic instead of a vibe
- [The Batch — DeepLearning.AI newsletter](https://www.deeplearning.ai/the-batch/) — Subscribe. The weekly signal-over-noise habit that keeps your inflection-point picture current.
  - Set a weekly signal-over-noise habit to stay current
  - Keep your inflection-point picture from going stale
  - Filter AI news down to what a leader actually needs
- [a16z — State of AI](https://a16z.com/state-of-ai/) — The investor lens: AI-native economics, valuations, and business-model shifts.
  - See AI-native economics and business-model shifts through an investor lens
  - Anticipate where valuations and models are moving
  - Read the investor case without adopting its incentives
- [Your own industry — 3 agentic case studies (research task)](https://www.google.com/search?q=agentic+AI+case+study+enterprise) — Find three companies already leveraging agentic approaches in or near your industry. Primary sources over hype pieces.
  - Find three agentic case studies in or near your own industry
  - Prefer primary sources over vendor testimonials
  - Judge what is actually shipping against what is promised

#### Project — Industry Disruption Analysis

Choose one industry you know deeply. Write a 2-page analysis: "How the AI inflection point is already disrupting (or will disrupt) this industry in the next 3 years" — covering cost collapse, capability explosion, and the time-to-X collapse. Document 3 companies using agentic approaches today, and build a simple AI Inflection Timeline (2022–2027) of the capability jumps that matter to this industry.

**Deliverable:** portfolio/w02-inflection-analysis.md — the 2-page analysis, 3 company snapshots, and the timeline.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Economic argument | 30% | Uses real numbers (cost per token/task trends, adoption data) rather than vibes; distinguishes efficiency gains from business-model change. |
| Agentic vs previous waves | 20% | Explains concretely why this differs from RPA/traditional ML for THIS industry — what becomes possible, not just cheaper. |
| Company evidence | 25% | Three real companies with sourced descriptions of their agentic approach and observable results; no press-release-only examples. |
| Timeline & judgment | 25% | Timeline picks capability jumps relevant to the industry, and the analysis takes a position — where value moves, who is exposed, what a leader should do in the next 2 quarters. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. The two simultaneous forces defining the current inflection point are…**

   a. Rising costs and shrinking capability
   b. Falling cost, rising capability
   c. More regulation and much less compute
   d. Fewer models but many more vendors

**2. "Time-to-X collapse" refers to…**

   a. Servers simply responding faster
   b. The shrinking idea-to-output gap
   c. Noticeably shorter overall working hours
   d. Faster model training runs overall

**3. Agentic AI differs from RPA fundamentally because…**

   a. It is cheaper to license
   b. RPA replays fixed clicks; agents adapt
   c. The agents only ever work up in the cloud
   d. RPA is newer

**4. "AI-native economics" most refers to…**

   a. Buying GPUs in bulk
   b. Business models built assuming near-zero marginal cost of cognition — e.g. revenue per employee outliers
   c. Charging more for AI features
   d. Paying engineers in equity

**5. Why is "waiting for the next budget cycle" newly dangerous at an inflection point?**

   a. Budgets are illegal to delay
   b. Capability compounds while rivals accelerate
   c. The vendors always raise their prices annually
   d. Auditors disapprove

**6. The most defensible way to size AI's impact on YOUR industry is…**

   a. Quoting the biggest consulting number
   b. Mapping specific value-chain steps to capabilities
   c. Simply waiting for a peer to publish their results first
   d. Running an employee survey

**7. A "capability jump" worth putting on your timeline is one that…**

   a. Made headlines
   b. Crossed a reliability/cost threshold that unlocks a real workflow in your industry
   c. Won a benchmark
   d. Was demoed at a conference

**8. Foundation models differ from traditional ML projects in that they…**

   a. They need more labelled data per task
   b. General-purpose, adapted per task
   c. They only ever work for spoken language
   d. Are always cheaper

**9. Which observation most strongly signals genuine agentic adoption (vs AI theater) at a company?**

   a. An AI mention buried in the annual report
   b. A chatbot on the website
   c. Workflows redesigned around agents
   d. An internal prompt-writing training course

**10. The strategic risk of over-indexing on ONE flashy AI demo is…**

   a. Demos are illegal to share
   b. Confusing a capability spike in a narrow case with reliable performance across your real distribution of work
   c. Demos expire
   d. Vendors dislike it

**11. Revenue per employee is a useful inflection metric because…**

   a. The HR team already carefully tracks it
   b. Output decouples from headcount
   c. It correlates closely with office size
   d. Regulators require it

**12. Your industry analysis should end with…**

   a. A neutral summary of all the various views
   b. A position on where value moves
   c. A long list of the possible vendors
   d. A disclaimer

### Module 3 — The Leap to Agentic AI: From Generation to Action

**Outcome:** Distinguish generative AI, automation, and true agents — and ship your first working agentic workflow.

**Leadership lens:** The difference between "AI that answers" and "AI that acts" is the difference between a research analyst and a delegated employee. Leaders must know first-hand what delegation to software feels like — including the loss-of-control moment.

**Apply-at-work mission — Automate one monitoring task:** Pick one thing you or your team checks manually (inbox, dashboard, RSS, sheet) and build an n8n/Make agentic workflow that monitors it and takes an action (summarise + notify). Run it for real.

**Reflection:** Where did I feel a loss of control this week when the agent acted for me? Which controls restored my confidence — and which were just comfort theatre?

#### Resources

- [n8n — docs & AI agent nodes](https://docs.n8n.io/) — Primary tool this week: visual workflow automation with real AI-agent nodes; free self-hosted.
  - Ship your first working automation with real AI-agent nodes
  - See a visual workflow do something end to end
  - Reach for a free self-hosted tool to prototype
- [CrewAI — documentation & quickstart](https://docs.crewai.com/) — Open-source multi-agent framework; run the research-agent example to see roles, tasks, and collaboration.
  - Run a multi-agent example and see roles, tasks and collaboration
  - Recognise when multiple agents beat a single one
  - Read the shape of role-based agent teams
- [LangGraph — agent loop tutorial](https://langchain-ai.github.io/langgraph/) — The clearest picture of the agent loop: plan → act → observe → repeat. Read for concepts even if you skip the code.
  - Picture the agent loop: plan → act → observe → repeat
  - Explain that loop on a whiteboard without the code
  - Tell a genuine agent from a scripted workflow
- [Anthropic — Building Effective Agents](https://www.anthropic.com/research/building-effective-agents) — The workflow-vs-agent distinction every executive should be able to make on a whiteboard.
  - Make the workflow-vs-agent distinction any executive should hold
  - Decide when a task genuinely needs an agent and when it does not
  - Name the guardrail a live agent needs
- [Building AI agents with n8n (curated tutorials)](https://www.youtube.com/results?search_query=n8n+ai+agent+tutorial) — Pick one end-to-end build video and follow along for your project.
  - Follow one end-to-end build for your own project
  - Turn a concept into a running prototype
  - Learn by shipping, not only watching

#### Project — First Agentic Workflow (Capstone Milestone 1)

Build your first real agentic workflow in n8n (or Make): an agent that monitors a source (email, RSS, form, sheet) and takes an action (summarise + notify, triage + route). Separately, run CrewAI's research-agent example and observe multi-agent collaboration. Then write a one-page comparison: what makes your workflow "agentic" vs a traditional Zapier-style automation — components (planner, memory, tools, executor, feedback) named explicitly. 🎯 This completes Capstone Milestone 1: a working agentic workflow you built yourself.

**Deliverable:** portfolio/w03-first-agent.md — workflow export/screenshots, what it does, the components map, and the automation-vs-agent comparison.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| It actually runs | 30% | The workflow executed on real inputs at least 5 times; you can show a run log and one output that reached you (or a colleague). |
| Component literacy | 25% | You can point at your workflow and name planner/memory/tools/executor/feedback loop — and honestly note which are missing. |
| Agentic vs automation analysis | 25% | The comparison identifies precisely where the LLM makes decisions vs where paths are fixed — no hand-waving about "smart automation". |
| Multi-agent observation | 20% | Notes from the CrewAI run: how agents divided the work, where they duplicated effort or drifted, and one thing that surprised you. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. The critical difference between generative AI and agentic AI is…**

   a. Model size
   b. Generative AI produces content; agentic AI acts
   c. Generative AI is a much older technology overall
   d. Agentic AI is always multi-agent

**2. The five core building blocks of an agent are…**

   a. GPU, CPU, RAM, disk, and the network
   b. Planner, memory, tools, executor
   c. Prompt, response, rating, retry, cache
   d. Input, output, database, UI, and API

**3. A Zapier automation that always runs the same fixed steps is…**

   a. Actually a genuinely true agent
   b. A fixed workflow, not an agent
   c. A full multi-agent system of some kind
   d. Simply an LLM running on its own

**4. Why do agents need a feedback/observation step in their loop?**

   a. To bill more tokens
   b. To check results and re-plan when reality diverges
   c. Purely for logging and regulatory compliance purposes
   d. To slow them down safely

**5. "Tools" in agent terminology means…**

   a. Various developer IDEs and code editors
   b. Functions/APIs the agent can call
   c. Assorted physical hardware accessories
   d. Prompt templates

**6. The safest first autonomy setting for a new workplace agent that can send messages is…**

   a. Full autonomy, logs reviewed monthly
   b. Draft-and-approve first
   c. Read-only mode for ever more
   d. No logging at all, to keep it simple

**7. In your n8n build, where does the "agentic" part actually live?**

   a. In the trigger node
   b. In the LLM node(s) deciding how to categorise/summarise/route based on content
   c. In the notification step
   d. In the credentials store

**8. Multi-agent collaboration (CrewAI-style) is best justified when…**

   a. When a single prompt just gets far too long
   b. Distinct roles genuinely divide the work
   c. You simply want it to look more impressive
   d. Single agents are deprecated

**9. An agent monitoring your inbox misfiles an important email. The leadership-grade response is…**

   a. Shut down all agents
   b. Treat it as an incident: trace the decision, tighten the rule or gate, keep operating
   c. Ignore it as a one-off
   d. Blame the vendor

**10. Real enterprise agent examples today (calendar, research, support triage) share what property?**

   a. Complete full autonomy over the money
   b. Bounded domains with clear tools
   c. No memory at all
   d. They replace entire whole departments

**11. The "loss of control" feeling when an agent acts for you is best handled by…**

   a. Never delegating actions
   b. Designing observability and reversibility
   c. Only ever running the agents fully manually
   d. Trusting the vendor

**12. Why does this programme make executives BUILD a workflow rather than just study agents?**

   a. To convert them into engineers
   b. First-hand delegation-to-software experience changes how they lead, buy, and govern it
   c. Because reading is ineffective
   d. To fill the week

---

## Phase 2 — Building: Vibe Coding & Agentic Workflows (weeks 4–8)

### Module 4 — Vibe Coding Fundamentals

**Outcome:** Build your first functional app with natural language — and develop judgment for when vibe coding shines vs when it bites.

**Leadership lens:** You don't need to become an engineer — you need to stop being blocked by the absence of one. A leader who can prototype an idea before the next steering meeting changes the tempo of the whole organisation.

**Apply-at-work mission — Ship a tool your team uses:** Build one small internal tool with Claude Artifacts / Cursor (dashboard, notes processor, checklist app) and put it in front of at least two colleagues. Capture their reaction and one improvement request.

**Reflection:** What did it feel like to build something without "knowing how to code"? Where did vibe coding fail me, and what did that teach me about reviewing AI work I can't fully verify?

#### Resources

- [Karpathy on vibe coding & AI-assisted development](https://www.youtube.com/results?search_query=karpathy+vibe+coding) — From the person who coined the term: what it is, why it works, where it ends.
  - Define vibe coding, why it works and where it ends — from the source
  - Set realistic expectations for natural-language building
  - Know the boundary between prototype and production
- [Cursor — official site & tutorials](https://cursor.com/) — The AI-first editor. Install it; you will describe, review, and refine rather than type code.
  - Describe, review and refine code rather than typing it
  - Build the judgment to review AI-written code
  - Ship a first functional tool without hand-coding
- [Anthropic — Claude Artifacts](https://www.anthropic.com/news/claude-artifacts) — The zero-setup build surface: describe an app, get a running one — ideal for your first tool.
  - Describe an app and get a running one with zero setup
  - Use the fastest surface for a first build
  - Prototype an idea in minutes to test it
- [Building real apps with AI assistants (curated)](https://www.youtube.com/results?search_query=building+apps+with+cursor+claude) — Watch one full honest build — including the parts where the AI gets it wrong.
  - Watch one honest build, including where the AI gets it wrong
  - Set expectations for the messy parts
  - Learn to spot where AI-assisted building breaks

#### Project — Ship an Internal Tool by Describing It

Using only natural-language prompting (Claude Artifacts or Cursor), build a small but genuinely useful internal tool: a personal dashboard, meeting-notes processor, decision log, or content repurposer. Put it in front of at least two colleagues. Then attempt to significantly improve an existing small script or spreadsheet process by vibe coding, and write a short honest critique: "Where vibe coding failed me and what I had to fix manually."

**Deliverable:** portfolio/w04-vibe-tool.md — what you built, prompts that mattered, colleague feedback, and the failure critique.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Usefulness | 30% | The tool addresses a real recurring annoyance; two colleagues used or reviewed it and one asked to keep it. |
| Prompt-to-product skill | 25% | You iterated: initial description → review → refinement cycles documented, showing you steered rather than accepted. |
| Critical judgment | 30% | The critique names specific failures (logic, edge cases, security assumptions) and what human judgment had to add — no cheerleading. |
| Risk awareness | 15% | You can state what this tool must NOT be used for (real data? external users?) and why. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. "Vibe coding" means…**

   a. Coding with music on
   b. Describing intent in natural language and letting AI generate the implementation, which you review and refine
   c. Copying from Stack Overflow
   d. Programming without tests

**2. The most important skill vibe coding does NOT remove is…**

   a. Sheer raw typing speed at the keyboard
   b. Validating it does what you intended
   c. Knowing all of the keyboard shortcuts
   d. Memorising syntax

**3. For a leader, the strategic value of being able to prototype is…**

   a. Completely replacing the whole engineering team
   b. Changing tempo — ideas before the next meeting
   c. Saving licence fees
   d. Winning hackathons

**4. Vibe-coded tools most dangerously accumulate…**

   a. A large amount of unnecessary disk usage over time
   b. Hidden technical debt: unexamined edge cases
   c. Too many features
   d. Licensing costs

**5. The right corporate policy for a useful vibe-coded prototype handling real customer data is…**

   a. Ship it — it works
   b. Treat it as a validated idea: hand to engineering for a hardened rebuild before real data touches it
   c. Keep it secret from IT
   d. Add a password

**6. When AI-generated code fails, it usually fails…**

   a. Very loudly and really quite obviously
   b. Plausibly — fine on the happy path
   c. Only ever at the compile-time stage itself
   d. Entirely at random and unpredictably

**7. Which request is vibe coding CURRENTLY best suited for?**

   a. A core banking system ledger
   b. A personal dashboard tool
   c. Medical-device firmware code
   d. A high-frequency trading engine

**8. The best prompting pattern for building beats one giant prompt because…**

   a. Very long prompts always end up costing a lot more money
   b. Iterative describe → inspect → refine catches drift
   c. Models refuse long prompts
   d. Short prompts look professional

**9. "I had to fix it manually" moments primarily teach leaders…**

   a. That AI is basically completely useless
   b. Where AI competence ends today
   c. That they should probably code a lot more
   d. Nothing useful

**10. A colleague proudly ships a vibe-coded customer-facing app with no engineering review. Your first question is…**

   a. Which particular model did you use for it?
   b. Malicious input, and who maintains it?
   c. How many separate prompts did it take you?
   d. Can I have the prompt?

**11. The "describing outcomes vs writing code" shift most resembles which leadership transition?**

   a. Engineer → manager
   b. Intern → junior analyst
   c. Employee → startup founder
   d. Manager → the board member

**12. Your vibe-coded tool works but you can't explain HOW. The leadership risk is…**

   a. None — results matter
   b. You cannot reason about its failure modes, so you cannot responsibly decide where it may be used
   c. Colleagues will ask questions
   d. It might be plagiarised

### Module 5 — Vibe Coding Advanced + API Integration

**Outcome:** Move beyond toys: build multi-step, API-connected applications that touch real business systems.

**Leadership lens:** The moment your prototype calls a real API, it stops being a demo and starts being a system — with error handling, credentials, and consequences. This is where executive prototypes earn (or lose) engineering's respect.

**Apply-at-work mission — Connect a prototype to a real system:** Extend a workflow or app so it calls at least one real external API (search, CRM, sheet, calendar) and processes the result with a second AI step. Document what you had to fix by hand.

**Reflection:** What architecture decisions did I make this week without realising they were architecture decisions? What would I now ask an engineering team that I couldn't have asked before?

#### Resources

- [Dify.ai — documentation](https://docs.dify.ai/) — Visual + code hybrid agent builder with API support — this week's main construction kit.
  - Build a multi-step, API-connected app with a visual+code hybrid
  - Move beyond toys to something touching real data
  - Choose a construction kit for a real integration
- [FlowiseAI — documentation](https://docs.flowiseai.com/) — Open-source visual agent builder; a strong alternative to Dify — skim to compare mental models.
  - Compare an open-source builder's mental model against Dify
  - Pick between builders on how they model agents
  - Know a strong alternative before you commit
- [OpenAI — function calling / tools guide](https://platform.openai.com/docs/guides/function-calling) — How models call external tools reliably. The conceptual core of every integration you'll build.
  - Explain how a model calls external tools reliably
  - Grasp the conceptual core of every integration you'll build
  - Design a tool call that actually works
- [Anthropic — tool use documentation](https://docs.anthropic.com/en/docs/build-with-claude/tool-use) — Same concept, Claude flavour: schemas, when the model decides to call, and error handling.
  - Define tool schemas and when the model decides to call
  - Handle tool errors inside an integration
  - Apply the same concept in Claude's flavour

#### Project — Multi-Step Research Agent with Real APIs

Build a multi-step agent that: (1) takes an input, (2) calls at least one real external API (web search, CRM, sheets, calendar), (3) processes results with a second AI step, and (4) outputs a structured action or report. The classic version: a personal research agent that searches the web and synthesises findings into a structured brief. Document your architecture decisions and everything you corrected by hand.

**Deliverable:** portfolio/w05-api-agent.md — architecture sketch, the working flow (screenshots/export), sample outputs, and the corrections log.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| End-to-end flow | 30% | Input → API call → AI processing → structured output all work on 5+ real runs, including one where the API returned something unexpected. |
| Architecture articulation | 25% | A simple diagram + prose naming each step, what can fail there, and what happens when it does. |
| Structured output | 20% | The final output follows a consistent, consumable format (fields, sections) — something a downstream system or colleague could rely on. |
| Corrections log | 25% | Honest record of what the AI got wrong and your fixes — evidence you reviewed rather than trusted. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. "Function calling" / tool use lets an LLM…**

   a. Rewrite its own weights
   b. Emit a structured request that your system executes (API call, query), returning results to the model
   c. Call phone numbers
   d. Compile code faster

**2. Why do multi-step agents need STRUCTURED outputs between steps?**

   a. They simply look a good deal tidier overall
   b. Downstream steps parse fields, not prose
   c. The regulators specifically require the use of JSON
   d. It reduces tokens

**3. An external API returns an error mid-flow. A well-designed agent…**

   a. Just crashes very visibly indeed
   b. Retries, falls back, reports
   c. Invents plausible data to continue
   d. Ignores the step

**4. The biggest NEW risk when your prototype connects to real business systems is…**

   a. Slower demos
   b. Consequences: real writes and real data exposure
   c. Noticeably higher overall token bills every month
   d. More complex prompts

**5. API credentials in an agent workflow should be…**

   a. Simply pasted into the prompt
   b. In the credential vault
   c. Shared in team chat for convenience
   d. Hard-coded but lightly obfuscated

**6. "Least privilege" for an agent's tools means…**

   a. Simply using the very cheapest available API tier
   b. Granting only the permissions the task needs
   c. Running fewer steps
   d. One tool per agent

**7. Chaining two AI steps (process API results with a second call) is useful because…**

   a. Two separate models are pretty much always much smarter
   b. Each step gets a focused job, improving reliability
   c. It doubles the billing
   d. Single calls are deprecated

**8. Your research agent cites a "fact" not present in any retrieved source. This is…**

   a. Actually quite a useful bonus insight
   b. A grounding failure in synthesis
   c. Entirely normal and perfectly fine
   d. An API bug

**9. Rate limits (429 responses) from an API should be handled by…**

   a. Failing the entire whole run
   b. Backing off and retrying
   c. Switching to other vendors immediately
   d. Simply ignoring them entirely

**10. Documenting architecture decisions matters for an executive builder because…**

   a. It pads the portfolio
   b. It converts building experience into the vocabulary you'll use to govern engineering teams and vendors
   c. Engineers demand it
   d. It is required for the certificate

**11. Which design makes an agent's work AUDITABLE?**

   a. Deleting the logs to save some space
   b. Logging every step fully
   c. Using a much bigger, better model
   d. Password-protecting the whole workflow

**12. The corrections log ("what I fixed by hand") is required because…**

   a. Everyone makes mistakes
   b. It builds your calibration of where AI-generated systems need human verification — the core delegation skill
   c. It fills the report
   d. It shames the model

### Module 6 — Designing Workflows for Agentic Systems

**Outcome:** Think computationally: map real business processes into agent-friendly workflows with clear decisions, exceptions, and human gates.

**Leadership lens:** Process mapping is old; mapping for agents is new. Every fuzzy handoff a human papers over becomes a failure mode when an agent runs it. The canvas you draw this week is the leadership artefact of the whole programme.

**Apply-at-work mission — Canvas a real workflow with a colleague:** Take one painful manual process from your organisation and complete the Agentic Workflow Canvas with the person who runs it: inputs, decisions, tools, outputs, exceptions, and the 5 places human judgment stays.

**Reflection:** Which steps of "my" processes could I actually not describe precisely when forced to? What does that say about how much of my organisation runs on tacit knowledge?

#### Resources

- [Computational Thinking for Problem Solving (Penn / Coursera)](https://www.coursera.org/learn/computational-thinking-problem-solving) — Free to audit. Decomposition, abstraction, patterns — the thinking layer under every workflow you'll design.
  - Apply decomposition, abstraction and patterns to a real process
  - Build the thinking layer under every workflow you design
  - Break a business process into agent-friendly steps
- [n8n — workflow design best practices](https://docs.n8n.io/workflows/) — How practitioners structure real workflows: naming, error paths, sub-workflows.
  - Structure a workflow with naming, error paths and sub-workflows
  - Design for failure, not just the happy path
  - Make a workflow maintainable by someone else
- [Multi-agent system design patterns (arXiv)](https://arxiv.org/search/?query=multi+agent+system+design+patterns) — Skim 1–2 recent surveys: sequential, hierarchical, swarm — patterns you'll map onto team structures.
  - Distinguish sequential, hierarchical and swarm patterns
  - Map an agent pattern onto a team structure
  - Pick a pattern that fits the process
- [Excalidraw](https://excalidraw.com/) — Your canvas tool this week — fast, free diagramming for the workflow maps you'll produce.
  - Diagram a workflow map fast and free
  - Make an agent design legible to stakeholders
  - Produce the visual your project needs this week

#### Project — Agentic Workflow Canvas (Capstone Milestone 2)

Take one painful manual process from your work or life and map it completely as an agentic workflow: inputs, decision points, tools, outputs, exception paths — and the 5 places human judgment must remain. Then design (on paper) a multi-agent system for a content or analysis pipeline (e.g. Researcher + Writer + Editor + Publisher) with a visual diagram and structured specification. 🎯 This completes Capstone Milestone 2: a real process, fully mapped for agents.

**Deliverable:** portfolio/w06-workflow-canvas.md + diagram — the completed canvas, the multi-agent spec, and the 5 human-judgment points.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Process fidelity | 25% | Mapped with the person who actually runs the process (or your honest first-hand knowledge); includes the messy exceptions, not the idealised flow. |
| Decision & exception design | 30% | Every decision point has defined criteria; every exception has a path (retry, fallback, escalate) — nothing ends in "somehow". |
| Human-judgment placement | 25% | The 5 human gates are placed where stakes or ambiguity are genuinely high, with rationale — not sprinkled for comfort. |
| Multi-agent spec quality | 20% | Roles have distinct responsibilities and tools; handoffs specify what artifact passes between agents in what format. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. The three levels of workflow abstraction this week are…**

   a. Code, config, and the docs
   b. Task → process → system
   c. Input, output, and storage
   d. Team, department, company

**2. Computational thinking, for a non-engineer leader, chiefly means…**

   a. Learning Python
   b. Decomposing fuzzy work into explicit steps, decisions, and data — precisely enough to delegate
   c. Doing mental arithmetic
   d. Thinking like a computer

**3. Why do human workflows break when handed directly to agents?**

   a. Agents are slow
   b. Human processes run on tacit knowledge
   c. The agents flatly refuse to do repetitive work
   d. APIs are unreliable

**4. A decision point in an agent-ready workflow map must have…**

   a. A single responsible senior executive
   b. Explicit criteria for each branch
   c. A hard timeout on the whole thing
   d. An AI model assigned

**5. Exception paths deserve MORE design attention than happy paths because…**

   a. They occur more often
   b. Failures cluster there, corrupting outcomes silently
   c. The auditors always read carefully through them first
   d. They are easier to design

**6. Human-in-the-loop gates are best placed where…**

   a. Wherever the work is at its most boring
   b. Stakes and ambiguity are high
   c. Wherever the agent happens to be fastest
   d. Wherever managers want more visibility

**7. Designing for observability means…**

   a. Making all the dashboards look pretty
   b. Each step can be inspected later
   c. Watching all the agents in real time
   d. Weekly status meetings

**8. In multi-agent design, a "hierarchical" pattern means…**

   a. The agents all vote on the decisions
   b. A supervisor delegates to workers
   c. All the agents just run in sequence
   d. All agents share one prompt

**9. The most common failure of first-time workflow mappers is…**

   a. Too many diagrams
   b. Mapping the idealised process, not the real one
   c. Simply using entirely the wrong tool for the job
   d. Too much detail

**10. A good handoff spec between two agents defines…**

   a. Their respective model versions
   b. The artifact and its format
   c. Their respective token budgets
   d. Which of them is more important

**11. Why produce the canvas BEFORE building anything?**

   a. Tools are expensive
   b. Design errors cost minutes on paper and weeks in production; the canvas is also your alignment artifact with stakeholders
   c. Building first is forbidden
   d. Diagrams impress leadership

**12. You cannot precisely describe a step in "your" process when forced to. The leadership lesson is…**

   a. The process is fine as folklore
   b. Your organisation runs on undocumented tacit knowledge — a risk AND the first thing agentification exposes
   c. Someone else should map it
   d. Skip that step

### Module 7 — Automating & Orchestrating Workflows

**Outcome:** Orchestrate multiple agents and tools into a production-style system with logging, monitoring, and a runbook.

**Leadership lens:** One agent is a trick; an orchestrated system is an operating model. Sequential, parallel, hierarchical, swarm — these patterns are org charts for software teammates, and you're the one drawing them.

**Apply-at-work mission — Stand up a multi-agent system:** Build a multi-agent workflow that solves a real problem end-to-end (e.g. research → qualify → draft outreach), add basic logging, and write a one-page runbook: how it works, what can go wrong, who to call.

**Reflection:** If this system ran while I slept, what's the worst thing it could plausibly do? Did my design catch that — or did I only design for the happy path?

#### Resources

- [LangGraph — build reliable agents (official tutorials)](https://langchain-ai.github.io/langgraph/tutorials/) — Code-level orchestration: graphs, state, and control over agent loops. Read for architecture even if you stay no-code.
  - Control agent loops with graphs and explicit state
  - Read the architecture of a reliable orchestration
  - Design control over an agent rather than hoping
- [CrewAI — multi-agent collaboration examples](https://github.com/crewaiinc/crewAI-examples) — Working examples of role-based agent teams — clone one close to your project and dissect it.
  - Clone and dissect a role-based agent team
  - Adapt a working example to your own project
  - See production-shaped multi-agent collaboration
- [n8n — advanced AI agent builds (curated)](https://www.youtube.com/results?search_query=n8n+advanced+ai+agent) — Full builds with memory, tools, and error handling — the production-shaped version of week 3.
  - Add memory, tools and error handling to a build
  - Turn a Week-3 prototype into something production-shaped
  - Handle the failure paths a demo skips
- [Dify.ai — agent workflow examples](https://docs.dify.ai/) — The visual+code hybrid path to orchestration; compare its trade-offs against n8n and CrewAI.
  - Compare the visual+code path against n8n and CrewAI
  - Weigh orchestration trade-offs across tools
  - Choose an orchestration approach on real criteria

#### Project — Orchestrated Multi-Agent System with a Runbook

Build a complete multi-agent system solving a real problem end-to-end — the reference build: automated lead/topic research → qualification/analysis → prepared output (outreach draft, brief, or report). Implement basic logging/observability so you can reconstruct any run. Then write a real runbook: how it works, what can go wrong, how to tell, and what to do about it.

**Deliverable:** portfolio/w07-orchestrated-system.md — system description + diagram, run logs from 5+ real executions, and the runbook.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| End-to-end operation | 30% | The full chain runs on real inputs without manual patching between steps; at least 5 logged runs including one failure you can explain. |
| Orchestration pattern choice | 20% | You chose sequential/parallel/hierarchical deliberately and can defend why against one alternative. |
| Observability | 25% | Logs let you answer "what did it do and why" for any run without guessing; one failure diagnosed FROM the logs. |
| Runbook quality | 25% | A colleague could operate the system from the runbook alone: failure symptoms → diagnosis → response, including the kill switch. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. Piloting vs productionising an agentic workflow differ mainly in…**

   a. The specific underlying model that gets used
   b. Error handling, monitoring, and ownership
   c. The prompt length
   d. The demo audience

**2. Sequential orchestration is the right default when…**

   a. You want the maximum possible speed
   b. Each step needs the last
   c. The agents disagree quite often
   d. You happen to have many GPUs

**3. A "swarm" pattern (many agents, loose coordination) is…**

   a. Always the most advanced choice
   b. Rarely justified — high coordination cost and failure surface for most business problems
   c. Required for production
   d. Cheaper than one agent

**4. The purpose of logging every agent step is…**

   a. Simply mere compliance box-ticking
   b. Reconstructing any run fully
   c. Slowing the whole system down safely
   d. Training data collection

**5. A runbook exists so that…**

   a. You remember your own system
   b. Someone who isn't the builder can operate it
   c. The auditors have some useful reading material
   d. The project looks professional

**6. The single most important control in any autonomous workflow is…**

   a. A much bigger model
   b. A fast kill switch
   c. A good many more agents
   d. Neatly colour-coded logs

**7. Two of your agents produce contradictory outputs mid-pipeline. Good orchestration design…**

   a. It simply just picks one of them at random
   b. Routes the conflict to a resolver
   c. It simply averages out all the answers
   d. Restarts everything

**8. Cost runaway in multi-agent systems typically comes from…**

   a. Expensive software licences
   b. Uncapped loops and retries
   c. Overly long variable naming
   d. Excessive logging overhead throughout

**9. "It worked in my demo" fails in operation most often because…**

   a. The demos simply use much better hardware
   b. Real inputs are messier and adversarial
   c. Users are less intelligent
   d. Networks are slower

**10. Basic observability for a business-owned agent system minimally includes…**

   a. The current GPU core temperature
   b. Status, traces, errors, cost
   c. Overall employee satisfaction levels
   d. Model weights

**11. The leadership reason to build this system yourself (once) is…**

   a. To replace your ops team
   b. To internalise the true effort, failure modes, and controls — calibrating how you buy, staff, and govern later
   c. To earn engineering credentials
   d. Because vendors are untrustworthy

**12. "What's the worst thing this system could plausibly do while I sleep?" is a design question because…**

   a. It motivates the team
   b. Autonomy means consequences happen without you present — so worst cases must be bounded by design, not attention
   c. It sounds impressive
   d. Regulations require insomnia

### Module 8 — Scaling from Pilot to SOP

**Outcome:** Know why most AI pilots die — and design the path from experiment to pilot to operational to scaled.

**Leadership lens:** Every organisation has a pilot graveyard. The 4-level maturity model, redesigned SOPs, and human+agent KPIs are how you become the leader whose pilots ship instead of stall.

**Apply-at-work mission — Maturity-assess a real pilot:** Run the Maturity Assessor on one AI pilot in your organisation (yours or someone else's). Produce a one-page scaling plan: current level, blockers to the next level, redesigned SOP, and two human+agent KPIs.

**Reflection:** Which of my organisation's KPIs actively punish agentic ways of working? What would I measure instead if I trusted the system?

#### Resources

- [McKinsey — why AI pilots fail to scale](https://www.mckinsey.com/capabilities/quantumblack/our-insights) — The canonical failure statistics and reasons — read one recent piece critically.
  - Cite why most AI pilots die, critically
  - Diagnose the failure pattern before you repeat it
  - Separate the real causes from the convenient ones
- [HBR — scaling AI in the enterprise (collection)](https://hbr.org/topic/artificial-intelligence) — Pick 2 articles on scaling/operating-model change; note what they say about KPIs and ownership.
  - Note what scaling takes: KPIs, ownership, operating-model change
  - Read two views on scaling and compare them
  - Plan for the operating-model shift, not just the tech
- [Agentic AI maturity models (arXiv + industry)](https://arxiv.org/search/?query=agentic+ai+maturity+model) — Emerging frameworks for experiment → pilot → operational → scaled; compare against this week's 4-level model.
  - Compare experiment → pilot → operational → scaled frames
  - Place your organisation on a maturity curve honestly
  - Design the path to the next level

#### Project — Pilot-to-Scale Plan (Capstone Milestone 3)

Take one of your own workflows (week 3 or 7) or a real pilot from your organisation and build its scaling plan on the 4-level maturity model: Experiment → Pilot → Operational → Scaled. Redesign one existing SOP to include agentic components, and define KPIs that measure BOTH efficiency and quality/judgment in the human+agent process. 🎯 This completes Capstone Milestone 3: a credible path from pilot to operations.

**Deliverable:** portfolio/w08-scaling-plan.md — maturity assessment, blockers per level, the redesigned SOP, and the new KPI set.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Honest maturity assessment | 25% | Current level justified with evidence; the temptation to grade your own pilot "operational" resisted. |
| Blocker analysis | 25% | Blockers to the next level are specific (ownership, error rate, integration, trust) with a named countermeasure each — not "needs more buy-in". |
| SOP redesign | 25% | The rewritten SOP specifies what the agent does, what humans do, and how exceptions and handoffs work — usable by the team tomorrow. |
| KPI design | 25% | At least two KPIs capture quality/judgment (not just speed/volume), each with a measurement method that would survive a sceptical CFO. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. The 4-level agentic maturity model runs…**

   a. Idea → to demo → to launch → to exit
   b. Experiment → Pilot → Operational
   c. Crawl → then walk → then run → then fly
   d. Dev → to test → to stage → to prod

**2. The most common reason AI pilots fail to scale is…**

   a. Weak models
   b. No operating-model change: unclear ownership, unchanged processes, missing integration — the org, not the tech
   c. Insufficient GPUs
   d. Prompt quality

**3. "Pilot purgatory" describes…**

   a. A whole series of failed experiments
   b. Pilots that never reach operations
   c. A set of quietly cancelled projects
   d. Vendor lock-in

**4. Moving from Pilot to Operational chiefly requires…**

   a. A bigger model
   b. Named ownership, error-handling, and integration
   c. Signing up a great many more pilot users than before
   d. A press release

**5. Traditional SOPs break in human+agent teams because…**

   a. The agents simply cannot read the documents at all
   b. They assume a human fills the gaps at each step
   c. SOPs are obsolete generally
   d. Compliance forbids agents

**6. A pure efficiency KPI ("tickets closed per hour") applied to an agentic process risks…**

   a. Nothing — efficiency is the goal
   b. Rewarding fast wrong answers
   c. Somewhat slower overall adoption
   d. Union complaints

**7. A good "quality/judgment" KPI for a human+agent workflow is…**

   a. Total agent runs
   b. Escalation appropriateness or correction rate
   c. The total number of separate prompts written out
   d. Tokens consumed

**8. Centralized vs federated workflow governance: the pragmatic answer for most orgs is…**

   a. Fully centralized — one AI team owns everything
   b. Central standards, federated build
   c. Fully federated — every single team alone
   d. External vendors entirely own all of it

**9. The "common failure pattern" of scaling a pilot everywhere immediately after one success is dangerous because…**

   a. Success is repeatable by default
   b. One context's success hides context-specific conditions; scaled rollouts meet different data, users, and edge cases
   c. It is too cheap
   d. Legal forbids it

**10. The right FUNDING model shift from pilot to scale is…**

   a. A one-off project budget forever
   b. From project to product funding
   c. No funding at all — it's automated now
   d. Charge users per prompt

**11. Your pilot's error rate is 8% with humans catching all errors. Before scaling, you must know…**

   a. Nothing at all — humans catch them
   b. Whether review holds at scale
   c. The model's exact parameter count
   d. The competitors' own error rates

**12. Redefining KPIs for agent-augmented teams is a LEADERSHIP task (not an analyst task) because…**

   a. Analysts are busy
   b. KPIs encode what the organisation values; changing them changes behaviour, careers, and incentives — political territory
   c. Leaders like dashboards
   d. It requires no data

---

## Phase 3 — Agentic AI Across the Business (weeks 9–11)

### Module 9 — Agentic AI in Operations & Sales

**Outcome:** Design agent systems for pipeline, accounts, and operations — with compliance, audit trails, and human oversight built in.

**Leadership lens:** Ops and sales are where agentic ROI shows up first and where compliance failures show up loudest. Lead research agents, follow-up sequences, approval workflows — leverage with a paper trail.

**Apply-at-work mission — Build a lead/ops research agent:** Build a "research & qualification agent" for a real ops or sales motion: input a company/case, output a structured brief. Include one human approval gate and show it to whoever owns that pipeline.

**Reflection:** Where is the line in my business between "agent prepares, human decides" and "agent decides"? Who should own moving that line — and is it currently owned by anyone?

#### Resources

- [n8n — CRM & business integrations](https://docs.n8n.io/integrations/) — HubSpot, Salesforce, sheets, mail — the connectors your ops/sales agents will live on.
  - Wire an ops or sales agent to HubSpot, Salesforce, sheets and mail
  - Know the connectors your agents will live on
  - Ground a pipeline agent in real business systems
- [Sales AI agents — practices & examples (curated)](https://www.youtube.com/results?search_query=sales+ai+agent+examples) — Watch 1–2 real implementations; note where humans stay in the loop and where compliance shows up.
  - Note where humans stay in the loop and where compliance shows up
  - Watch a real implementation critically
  - Design a sales agent with the right human gates
- [Compliant AI workflows in regulated settings (arXiv)](https://arxiv.org/search/?query=compliant+ai+workflows) — Skim one paper/report: audit trails, approval gates, and record-keeping for AI-assisted processes.
  - Build audit trails, approval gates and record-keeping into a process
  - Meet regulated-setting requirements for AI-assisted work
  - Know what compliance demands before you ship

#### Project — Lead Research & Qualification Agent

Build a "Lead Research & Qualification Agent" for a real ops or sales motion: input a company name (or case), output a structured brief — ICP fit, key people, recent news, likely pain points. Add an automated follow-up-preparation step that respects compliance rules, and design a human-in-the-loop approval gate for any high-value action. Show it to whoever owns that pipeline.

**Deliverable:** portfolio/w09-sales-ops-agent.md — the working flow, 3 sample briefs, the compliance notes, and pipeline-owner feedback.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Brief quality | 30% | Structured, sourced, decision-ready briefs a real seller/operator would use — verified against what they already know about one account. |
| Human approval gate | 25% | High-value actions demonstrably cannot proceed without human release; the gate shows WHAT the human approves, not just a yes button. |
| Compliance thinking | 25% | Data sources, consent constraints, and record-keeping named explicitly; nothing scraped or stored that the business couldn't defend. |
| Stakeholder validation | 20% | Feedback from the pipeline owner captured honestly, including what they would NOT trust the agent with. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. Ops and sales are typically the first agentic AI beachhead because…**

   a. The salespeople all love technology
   b. High-volume, structured work
   c. They tend to have the biggest budgets
   d. Compliance is entirely absent there

**2. A lead-research agent's output should be structured (ICP fit, people, news, pains) because…**

   a. It looks professional
   b. Sellers act on decision-ready fields
   c. Free text is basically impossible to use
   d. CRMs demand JSON

**3. Automated outreach WITHOUT human review risks…**

   a. Only a few minor spelling errors
   b. Compliance and brand damage
   c. Somewhat slower overall pipelines
   d. Higher monthly API bills to pay

**4. An audit trail in a sales/ops agent context means…**

   a. A full backup copy of the entire CRM system
   b. A record of what was done and by whom
   c. All of the customer call recordings kept
   d. The prompt library

**5. The human approval gate for high-value actions should show the approver…**

   a. A progress bar
   b. The action, its evidence, and its consequences
   c. Just a simple single approve button for them to click
   d. The model temperature

**6. "Pipeline intelligence" agents create most value by…**

   a. Replacing account executives
   b. Compressing research and preparation time so humans spend their hours on judgment and relationships
   c. Sending more emails
   d. Automating discounts

**7. Data used by a research agent must be checked for…**

   a. Font compatibility
   b. Provenance and permission: is it public, licensed, or consented — and would you defend its use?
   c. File size
   d. Language

**8. The agent confidently reports a "recent funding round" that never happened. The systemic fix is…**

   a. Immediately fire the whole agent
   b. Ground claims in cited sources
   c. Simply lower the temperature setting
   d. Add a disclaimer

**9. Which task should REMAIN human in an agent-augmented sales process?**

   a. General company news gathering
   b. Final deal-strategy judgment
   c. Routine CRM field record updates
   d. General meeting scheduling work

**10. Rolling this agent to the whole sales team after one rep's success requires first…**

   a. Nothing at all — the success clearly proves it
   b. Checking the workflow fits other segments
   c. A bigger model
   d. A launch party

**11. The pipeline-owner's "I wouldn't trust it with X" feedback is valuable because…**

   a. It usefully identifies the sceptics
   b. It maps the trust boundary clearly
   c. It can be safely ignored after launch
   d. It fills the report

**12. Measuring this agent's success should centre on…**

   a. Number of briefs generated
   b. Downstream outcomes: qualified-meeting rate, time-to-first-touch, and correction rate of briefs
   c. Tokens used
   d. Rep satisfaction only

### Module 10 — Agentic AI in Marketing & Customer Experience

**Outcome:** Build CX and content systems that scale personalisation without sacrificing brand, tone, empathy, or privacy.

**Leadership lens:** Marketing is the easiest place to deploy agents and the easiest place to damage a brand at machine speed. Brand-voice grounding, escalation paths, and consent-aware personalisation are leadership controls, not settings.

**Apply-at-work mission — Repurpose content with brand guardrails:** Build a content repurposing agent grounded in a real brand-voice document (yours or your employer's): one long-form piece → three platform-specific versions. Have the brand owner grade the outputs.

**Reflection:** What parts of my customers' experience should never be synthetic, even if no one could tell? Where do I draw that line and why?

#### Resources

- [Marketing AI agents & automation (curated)](https://www.youtube.com/results?search_query=marketing+ai+agents+automation) — Real content/CX automations; watch for how (or whether) they protect brand voice.
  - Watch how automations protect — or fail to protect — brand voice
  - Spot where personalisation scales and where it breaks
  - Design CX automation that keeps the brand intact
- [Anthropic — research & brand-safe AI systems](https://www.anthropic.com/research) — Grounding, steerability, and constitutional approaches — the technical roots of tone control.
  - Understand grounding and steerability as the roots of tone control
  - Explain the technical basis of brand-safe AI
  - Set expectations for controlling a model's voice
- [HBR — customer experience & AI case studies](https://hbr.org/topic/artificial-intelligence) — Pick one CX case study; note where personalisation crossed into creepy or synthetic.
  - Note where personalisation crossed into creepy or synthetic
  - Judge a CX case on the human line, not the tech
  - Draw your own boundary for acceptable personalisation
- [Dify / Flowise — support-agent templates](https://docs.dify.ai/) — Build surface for the support triage agent — templates get you to a working baseline fast.
  - Get a support-triage agent to a working baseline fast
  - Start from a template rather than a blank canvas
  - Ship a CX prototype this week

#### Project — Brand-Safe Content & CX Agents

Two builds. (1) A content repurposing agent grounded in a real brand-voice document: one long-form piece in, three platform-specific versions out — graded by the brand owner. (2) A customer-support triage agent design (build if time allows): categorise incoming requests, draft responses for routine cases, and define clear escalation paths to humans. Plus: a one-page personalisation policy covering privacy, consent, and tone.

**Deliverable:** portfolio/w10-brand-cx-agents.md — repurposer outputs + brand-owner grades, the triage design, and the personalisation policy.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Brand-voice fidelity | 30% | The brand owner grades outputs ≥7/10 on voice; deviations analysed — you know WHY it drifted where it did. |
| Platform adaptation | 20% | The three versions genuinely differ in form and register for their platforms, not just in length. |
| Escalation design | 25% | Triage rules name the categories agents may answer, must draft-only, and must hand straight to humans — with the signals that trigger each. |
| Personalisation policy | 25% | Privacy/consent/tone rules concrete enough to enforce: what data may personalise what, and what must never be synthetic. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. The core tension in agentic marketing is…**

   a. Overall cost versus raw speed
   b. Scale vs brand consistency
   c. SEO versus paid online advertising
   d. Plain text versus rich video

**2. Grounding a content agent in a brand-voice document works because…**

   a. It shortens the prompts you need to write out
   b. It generates against explicit constraints
   c. It reduces cost
   d. Legal requires it

**3. "Human oversight" in agentic content creation is best implemented as…**

   a. Reading everything after publication
   b. Review gates by risk tier
   c. Trusting the prompt
   d. Weekly audits only

**4. A support triage agent should hand to a human IMMEDIATELY when…**

   a. The queue happens to be long
   b. It detects distress or risk
   c. The customer writes in all caps
   d. After about three or so exchanges

**5. Customer journey orchestration with agents means…**

   a. Simply sending more emails per customer
   b. Coordinating touchpoints coherently
   c. Automating all of the churn surveys sent
   d. A/B testing everything

**6. Consent-aware personalisation requires…**

   a. Personalising everything that is possible
   b. Only data provided for that purpose
   c. Anonymous background tracking of them
   d. Longer privacy policies

**7. The biggest CX risk of routine-response automation is…**

   a. Occasional minor grammar mistakes here
   b. Confidently wrong replies at scale
   c. Somewhat slower responses overall now
   d. Higher costs

**8. "What should never be synthetic" is a leadership question because…**

   a. Engineers can't decide fonts
   b. It draws the brand's authenticity line — a values decision with commercial consequences, not a technical setting
   c. Regulators publish the list
   d. It changes weekly

**9. Measuring a repurposing agent purely on output volume incentivises…**

   a. Genuinely better quality
   b. Off-brand content
   c. Much better research
   d. Nothing at all — volume is neutral

**10. Brand-owner grading of agent outputs (your project) establishes…**

   a. Who is boss
   b. A quality baseline and a feedback signal
   c. A useful published marketing case study for it
   d. Legal cover

**11. An agent personalises a message using data the customer never knowingly shared. Even if legal, this is…**

   a. A win for relevance
   b. A trust liability: perceived surveillance damages the relationship personalisation was meant to deepen
   c. Standard practice — fine
   d. A/B testable

**12. Escalation paths must be designed BEFORE deploying CX agents because…**

   a. Documentation looks good
   b. Under load, undefined escalation collapses into either everything-to-humans or nothing-to-humans
   c. Agents demand them
   d. They cannot be changed later

### Module 11 — Agentic AI in HR & Finance

**Outcome:** Apply agents to people and money decisions with bias mitigation, auditability, and accountability that would survive a regulator.

**Leadership lens:** HR and Finance are high-leverage AND high-risk: résumé screening, anomaly detection, performance prep. The design question is never "can the agent do it?" — it's "who is accountable when it does?"

**Apply-at-work mission — Design a review-gated HR/Finance workflow:** Design (build if you can) one HR or Finance agent workflow with explicit bias-mitigation steps and human review gates — e.g. screening assistant or spend-anomaly reporter. Write down its audit trail.

**Reflection:** If an agent-assisted decision about a person turned out wrong, could I explain the decision chain to that person's face? What would need to change so I could?

#### Resources

- [SHRM — AI in HR (topic hub)](https://www.shrm.org/topics-tools/topics/artificial-intelligence) — The practitioner view: where HR is deploying AI and which risks dominate the conversation.
  - See where HR is deploying AI and which risks dominate the conversation
  - Take the practitioner view, not the vendor view
  - Name the HR-specific risks before you build
- [Deloitte — AI in finance](https://www2.deloitte.com/global/en/pages/financial-services/articles/ai-in-finance.html) — Finance-function use cases: analysis, forecasting, anomaly detection — and the governance wrapper.
  - Map finance use cases: analysis, forecasting, anomaly detection
  - Note the governance wrapper each one needs
  - Apply agents to money decisions with the right controls
- [Auditable AI workflows (arXiv)](https://arxiv.org/search/?query=auditable+ai+workflows) — Skim one paper: what makes an AI-assisted decision reconstructable after the fact.
  - Make an AI-assisted decision reconstructable after the fact
  - Build the audit trail a people or money decision requires
  - Meet auditability before deploying to sensitive functions

#### Project — Accountable HR/Finance Agent Workflows

Design — and build what you can — two sensitive-domain workflows: (1) a résumé screening + initial outreach assistant with explicit bias-mitigation steps and human review gates; (2) a financial anomaly detection + reporting agent (n8n + sheets + LLM analysis works). For each, document the audit trail: what is recorded, who is accountable for each decision, and how a challenged decision would be reconstructed and explained.

**Deliverable:** portfolio/w11-hr-finance-agents.md — both designs, the bias-mitigation measures, audit-trail specs, and an accountability map.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Bias mitigation | 30% | Concrete measures (structured criteria before screening, blind fields, sampled human audits of rejections) — not a statement that bias is bad. |
| Human accountability | 25% | Every consequential decision has a named human owner; the agent recommends, a person decides — and the map shows it. |
| Audit trail design | 25% | A challenged decision (rejected candidate, flagged transaction) can be fully reconstructed: inputs, criteria, model output, human action. |
| Risk/leverage judgment | 20% | The write-up distinguishes where automation is high-leverage (drudgery) vs high-risk (judgment about people/money) with defensible placements. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. HR and Finance are "high-risk, high-leverage" domains for agents because…**

   a. They happen to have by far the very most data
   b. Huge drudgery, but errors touch livelihoods
   c. Budgets are centralised
   d. They adopt slowly

**2. The correct division of labour for résumé screening is…**

   a. The agent decides, human is informed
   b. Agent summarises; human decides
   c. Agent decides, with an appeal option
   d. Human does everything

**3. AI screening can AMPLIFY hiring bias because…**

   a. Models dislike certain fonts
   b. It learns patterns from historical decisions — including the biased ones — and applies them consistently at scale
   c. It reads too fast
   d. Candidates game it

**4. A meaningful bias-mitigation step is…**

   a. Telling the model to be fair
   b. Defining structured job-relevant criteria BEFORE screening and auditing a sample of rejections against them
   c. Using a larger model
   d. Removing all human review

**5. For a financial anomaly-detection agent, a false NEGATIVE (missed anomaly) vs false POSITIVE (false alarm) trade-off should be set by…**

   a. Simply the model's own built-in defaults
   b. Finance leadership's risk appetite
   c. Whatever minimises the alerts sent
   d. The vendor

**6. "Auditability" in these domains concretely means…**

   a. Annual audits happen
   b. Any decision can be reconstructed later
   c. The CFO has a nice live dashboard for it
   d. Logs exist somewhere

**7. An agent flags an employee expense as anomalous. The next step should be…**

   a. An automatic payroll deduction
   b. A human reviews it first
   c. An automated warning email sent
   d. Just silent logging of it

**8. Explaining an agent-assisted decision "to the person's face" is the right design test because…**

   a. It is emotionally satisfying
   b. It forces reconstructable, criteria-based decisions — if you can't explain it, you shouldn't operate it
   c. Lawyers recommend it
   d. It shortens meetings

**9. Which HR task is the SAFEST early agent deployment?**

   a. Making termination decisions
   b. Drafting role descriptions
   c. Setting performance ratings
   d. Setting people's actual salaries

**10. "The agent recommends, the human decides" fails in practice when…**

   a. The agents are simply too slow
   b. Humans rubber-stamp it
   c. The recommendations are too good
   d. The data is all well-formatted

**11. Tracking how often human reviewers DISAGREE with agent recommendations is valuable because…**

   a. It usefully ranks all of the employees
   b. Near-zero disagreement is a warning
   c. It quietly trains up the model further
   d. It fills dashboards

**12. Regulatory exposure for AI in HR/Finance (GDPR/AI-Act-style rules) most concerns…**

   a. Model size limits
   b. Automated decisions about individuals
   c. Open-source software licensing terms only
   d. Cloud regions only

---

## Phase 4 — Leading in the Agentic Age (weeks 12–15)

### Module 12 — Judgment in an Age of Abundance

**Outcome:** Intelligence is abundant; judgment is scarce. Learn to lead with WHY–WHAT–HOW when execution is nearly free.

**Leadership lens:** When AI can execute anything in minutes, the bottleneck moves to whoever decides what's worth executing. Decision velocity and learning velocity become the new leadership metrics — framing beats doing.

**Apply-at-work mission — Run a WHY–WHAT–HOW conversation:** Apply the WHY–WHAT–HOW framework to one live strategic question and use it to structure a real 15-minute conversation with your manager or team. Note where the framework changed the outcome.

**Reflection:** Write the first draft of your leadership manifesto: how will I lead differently in an age of abundant intelligence? Which of my current strengths become commodities?

#### Resources

- [HBR — leadership in the age of AI (collection)](https://hbr.org/topic/leadership) — Pick 2 pieces on decision-making and leadership under AI; read for frameworks, not comfort.
  - Take frameworks for decision-making under AI, not comfort
  - Read two pieces critically for what actually transfers
  - Lead with judgment where intelligence is abundant
- [Mustafa Suleyman — The Coming Wave (talks & summaries)](https://www.youtube.com/results?search_query=mustafa+suleyman+the+coming+wave) — The abundance argument from one of its architects: what happens when intelligence is no longer scarce.
  - Understand the abundance argument from one of its architects
  - Anticipate what happens when intelligence is no longer scarce
  - Position judgment as the scarce, valuable resource
- [Human judgment in AI-augmented organisations (arXiv)](https://arxiv.org/search/?query=human+judgment+ai+organizations) — Skim recent research on where human judgment retains advantage — and where we overestimate it.
  - Locate where human judgment retains a real advantage
  - Spot where we overestimate our own judgment
  - Apply WHY–WHAT–HOW where judgment matters most

#### Project — Leadership Manifesto + WHY–WHAT–HOW in Practice

Deep-thinking week. (1) Write your personal leadership manifesto: "How I will lead differently in an age of abundant intelligence" — using an AI as a Socratic sparring partner that challenges rather than agrees. (2) Apply the WHY–WHAT–HOW framework to one live strategic challenge in your organisation and document the result. (3) Identify 3 areas where FASTER AI execution would create problems without better human framing.

**Deliverable:** portfolio/w12-manifesto.md — the manifesto, the WHY–WHAT–HOW application, and the 3 speed-traps.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Manifesto depth | 30% | Names specific behaviours you will start/stop/change — commitments a colleague could hold you to, not aspirational fog. |
| Socratic process | 20% | Evidence the AI challenged you: at least two positions you revised under questioning, documented. |
| Framework application | 30% | The WHY–WHAT–HOW analysis of a real challenge shows the layers separated cleanly — and reveals something the default framing missed. |
| Speed-trap insight | 20% | The 3 areas are real and defensible: places where execution velocity would amplify a framing error. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. The "AI paradox" of execution states…**

   a. The AI simply cannot execute anything at all
   b. Direction is the binding constraint
   c. Execution no longer really matters much
   d. Paradoxes sell books

**2. In the WHY–WHAT–HOW framework, leaders should increasingly concentrate on…**

   a. HOW — implementation detail
   b. WHY and WHAT — purpose and framing
   c. None of them; all three equally forever
   d. Whichever is urgent

**3. "Decision velocity" as a leadership metric means…**

   a. Deciding everything instantly
   b. The speed from information to committed, high-quality decision — matched to the new pace of execution
   c. Number of decisions per week
   d. Delegating all decisions

**4. "Learning velocity" matters more than stock of knowledge now because…**

   a. Knowledge is expensive
   b. Capability and context shift so fast that the rate of updating beats the depth of any static expertise
   c. Universities are slow
   d. Memory fades

**5. Moving from "executor" to "framer and orchestrator" means…**

   a. Doing less work
   b. Defining problems, constraints, and quality bars
   c. Simply attending a good many more meetings than before
   d. Approving others' work

**6. A team ships a polished, wrong deliverable in record time. The root failure is most likely…**

   a. A team of overly lazy agents
   b. Bad framing of the WHAT
   c. Simply insufficient compute
   d. Just plain bad luck really

**7. Using AI as a SOCRATIC partner (vs an answer engine) means…**

   a. Simply asking rather harder questions of it
   b. Having it interrogate your assumptions
   c. Using two entirely separate models at once
   d. Asking twice

**8. Which decision still deserves SLOW deliberation despite fast execution being available?**

   a. Reformatting a whole report
   b. An irreversible commitment
   c. Choosing everyone's meeting times
   d. Renaming a single small file

**9. "Intelligence is abundant, judgment is scarce" implies organisations should now compete on…**

   a. Raw underlying model access rights
   b. Problem selection and framing
   c. Big shared prompt libraries kept
   d. Compute supply contracts signed

**10. The most dangerous leadership response to abundant execution capacity is…**

   a. Careful, deliberate prioritisation
   b. Filling capacity with more
   c. Training up the whole team properly
   d. Measuring outcomes

**11. A leadership manifesto beats vague intentions because…**

   a. It reads well
   b. Written, specific commitments enable self-accountability and let others call the gap between stated and lived
   c. HR requires it
   d. It is shareable

**12. Framing work ("what problem are we actually solving?") resists automation because…**

   a. The models are generally quite bad at using words
   b. It needs context, stakes, and values judgment
   c. It is low-status work
   d. It changes rarely

### Module 13 — The Three Roles That Scale: Orchestrator, Architect, Steward

**Outcome:** Master the three roles that gain value in agentic organisations — and diagnose which one you and your team are missing.

**Leadership lens:** Traditional management scales by headcount; agentic organisations scale by orchestration, architecture, and stewardship. Knowing which role you naturally play — and which your team lacks — is a career-defining insight.

**Apply-at-work mission — Role-map your team:** Map yourself and your team against Orchestrator / Architect / Steward. Identify the missing role, then draft a job description for an "Agentic AI Orchestrator" (or the role you lack) in your organisation's language.

**Reflection:** Which of the three roles do I gravitate to under pressure — and is that the role my organisation actually needs most from me right now?

#### Resources

- [McKinsey — future of work: new roles in AI organisations](https://www.mckinsey.com/featured-insights/future-of-work) — What roles grow as AI absorbs execution — evidence for the Orchestrator/Architect/Steward thesis.
  - Identify which roles grow as AI absorbs execution
  - Ground the Orchestrator/Architect/Steward thesis in evidence
  - Anticipate the roles your organisation will need
- [HBR — organisational design (collection)](https://hbr.org/topic/organizational-design) — One or two pieces on designing AI-augmented teams; note who holds authority and accountability.
  - Note who holds authority and accountability in AI-augmented teams
  - Design a team around the three scaling roles
  - Read org-design pieces for structure, not slogans
- [Stewardship & governance in agentic systems (arXiv)](https://arxiv.org/search/?query=ai+stewardship+governance) — The research frame for the Steward role: trust, oversight, and accountability at system level.
  - Frame the Steward role around trust, oversight and accountability
  - Understand system-level governance as a role, not a checkbox
  - Assign stewardship deliberately

#### Project — Role Map & Agentic Team Blueprint (Capstone Milestone 4)

Map your current role and team against the Orchestrator–Architect–Steward model: who covers what, where the gaps are, which role you gravitate to. Design a small team structure (3–5 people + agents) optimised around one primary role for a real objective. Write a serious job description for an "Agentic AI Orchestrator" (or your organisation's missing role) in your company's language. 🎯 This completes Capstone Milestone 4: an organisational design you could actually propose.

**Deliverable:** portfolio/w13-role-blueprint.md — the role map with gaps, team design, and the job description.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Honest role mapping | 25% | Real names/functions mapped with evidence; the missing role identified from observed failures, not theory. |
| Team design coherence | 30% | The 3–5 person + agents structure names each human's role, each agent's scope, and the decision rights between them. |
| Job description quality | 25% | The JD would survive HR review: responsibilities, competencies, success measures at 6/12 months — in your org's idiom. |
| Self-awareness | 20% | Your own gravitational role identified with supporting evidence, plus a development plan for your weakest of the three. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. The three roles that scale in agentic organisations are…**

   a. The coder, the tester, the deployer
   b. Orchestrator, Architect, Steward
   c. The manager, director, and the VP
   d. The analyst, designer, and marketer

**2. The Orchestrator's core skill is…**

   a. Writing prompts
   b. Aligning humans and agents across silos — sequencing, handoffs, and shared context toward an outcome
   c. Approving budgets
   d. Debugging code

**3. The Architect's core skill is…**

   a. Ongoing people and team-management work
   b. Designing workflows and decision logic
   c. Model training
   d. Vendor negotiation

**4. The Steward's core skill is…**

   a. Aggressive cost cutting throughout
   b. Ensuring trust and accountability
   c. Loudly marketing all of the AI wins
   d. Writing policies nobody reads

**5. These roles scale differently from traditional management because…**

   a. They generally tend to pay a lot more
   b. Leverage grows with the workflows
   c. They genuinely need no real skills
   d. They avoid meetings

**6. A team ships agent workflows fast but keeps having governance incidents. The missing role is…**

   a. Orchestrator
   b. Steward
   c. Architect
   d. More engineers

**7. Workflows exist and are governed, but nothing connects across departments and agents duplicate work. Missing role…**

   a. Steward
   b. Orchestrator
   c. Architect
   d. Project manager

**8. Everyone uses AI enthusiastically but workflows are ad-hoc, fragile, and undocumented. Missing role…**

   a. Orchestrator
   b. Architect
   c. Steward
   d. Trainer

**9. Middle management's traditional information-relay function is threatened because…**

   a. The executives all got much friendlier
   b. Agents relay information instantly
   c. Most offices simply went fully remote
   d. Budgets shrank

**10. A serious "Agentic AI Orchestrator" JD should measure success by…**

   a. Number of prompts written
   b. Cross-functional workflow outcomes: cycle time, quality, and adoption of orchestrated human+agent processes
   c. Meetings hosted
   d. Certifications earned

**11. Knowing which role YOU gravitate toward matters because…**

   a. It largely determines your salary
   b. You over-supply your own role
   c. The roles are all quite permanent
   d. The HR team always tends to ask

**12. A 4-person team + agents "optimised for one primary role" means…**

   a. Everyone does everything
   b. The team's structure, rituals, and metrics centre on that role's leverage — e.g. an architecture pod that designs workflows for many teams
   c. Three people are idle
   d. One person rules

### Module 14 — Cognitive Capital vs Cognitive Debt

**Outcome:** Design personal and team practices where AI amplifies thinking instead of replacing it — before automation bias compounds.

**Leadership lens:** Every convenience is a loan against a skill. Teams that offload thinking accrue cognitive debt that comes due at the worst moment — in a crisis, when the AI is wrong and nobody can tell. Leaders set the repayment schedule.

**Apply-at-work mission — Audit your AI usage:** Run the Cognitive Capital Audit on your own last two weeks of AI usage: where did AI deepen your thinking, where did it replace it? Adopt one "Thinking Amplification Protocol" and practice it for 5 days.

**Reflection:** Which thinking skill have I quietly stopped practicing since I started using AI daily? Do I want it back — and what's my plan if the answer is yes?

#### Resources

- [Automation bias research (Google Scholar)](https://scholar.google.com/scholar?q=automation+bias+ai) — The evidence base: humans defer to automated suggestions even against their own judgment.
  - Cite evidence that people defer to automation against their own judgment
  - Recognise automation bias in your own team
  - Design against reflexive deference
- [Cognitive offloading and AI (arXiv)](https://arxiv.org/search/?query=cognitive+offloading+ai) — What happens to skills we stop practicing — the cognitive-debt mechanism, measured.
  - Understand the cognitive-debt mechanism, as measured
  - Anticipate which skills atrophy when offloaded
  - Decide what thinking must stay in-house
- [Building metacognition with AI (curated)](https://www.youtube.com/results?search_query=metacognition+with+ai+learning) — Practical approaches to thinking-about-thinking with AI as amplifier rather than substitute.
  - Use AI to amplify thinking rather than replace it
  - Build thinking-about-thinking practices with AI
  - Design team habits that grow cognitive capital
- [NotebookLM / Obsidian — reflective knowledge work](https://notebooklm.google/) — Tools for the thinking-partner pattern: your notes + AI interrogation, not AI answers from nowhere.
  - Run the thinking-partner pattern: your notes plus AI interrogation
  - Avoid AI answers from nowhere by grounding in your own knowledge
  - Set up a reflective knowledge-work loop

#### Project — Cognitive Capital Audit + Thinking Amplification Protocol

Audit your own AI usage over the past two weeks: for each significant use, classify — did AI deepen your thinking (capital) or replace it (debt)? Identify 3 areas of accumulating cognitive debt. Design your personal "Thinking Amplification Protocol" — a workflow where AI deepens understanding instead of shortcutting it — and practice AI Socratic dialogue on one complex topic: the AI challenges your assumptions rather than answering.

**Deliverable:** portfolio/w14-cognitive-audit.md — the usage audit, 3 debt areas, your protocol, and the Socratic dialogue transcript with commentary.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Audit honesty | 30% | Real usage examined without flattering yourself; at least one uncomfortable finding about your own offloading. |
| Debt diagnosis | 20% | The 3 debt areas name specific skills degrading (structuring arguments, estimation, first-draft thinking) with evidence. |
| Protocol design | 30% | The protocol is concrete and practiced for 5 days: think first → AI challenges → revise — not an aspiration but a routine with a trigger. |
| Socratic transcript | 20% | The dialogue shows genuine challenge: assumptions surfaced, at least one position revised, commentary on what the process felt like. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. Automation bias is…**

   a. Simply preferring newer tools
   b. Over-trusting the machine
   c. An outright hatred of manual work
   d. A kind of model training error

**2. Cognitive debt accumulates when…**

   a. You think too hard
   b. Skills atrophy because AI does them for you
   c. Your working memory quietly fills right up now
   d. You take notes

**3. Cognitive capital, by contrast, is built when…**

   a. You avoid using AI entirely and completely
   b. AI use deepens your understanding
   c. You save prompts
   d. You read more

**4. The clearest warning sign of personal cognitive debt is…**

   a. Simply using AI every single day
   b. Being unable to judge the work
   c. Generally preferring dictation now
   d. Typing slower

**5. "Think first, then consult AI" beats "ask AI first" for important problems because…**

   a. It is more polite
   b. Forming your own position first preserves the thinking skill and makes you a critic rather than a consumer of the AI's answer
   c. AI answers improve later in the day
   d. It saves tokens

**6. For a TEAM, cognitive debt shows up as…**

   a. Consistently much longer meetings
   b. Nobody can explain the decisions
   c. A great deal more documentation produced
   d. Slower hiring

**7. Metacognition in this context means…**

   a. Thinking quickly
   b. Monitoring your own thinking closely
   c. Various memory and recall techniques used
   d. Meditation

**8. A Thinking Amplification Protocol should REQUIRE…**

   a. Avoiding AI for the hard problems
   b. Your own draft before AI
   c. Using two separate AI models
   d. Longer prompts

**9. Leaders bear special responsibility for cognitive debt because…**

   a. They personally use AI the very most
   b. They set the norms teams copy
   c. Their own skills matter far less
   d. They approve all of the tools used

**10. Which AI use pattern builds capital rather than debt for a strategy question?**

   a. "Write our 3-year strategy"
   b. "Here is my draft strategy and reasoning — find the three weakest assumptions and argue against them"
   c. "Summarise our industry"
   d. "What would McKinsey say?"

**11. The crisis-scenario argument for maintaining thinking skills is…**

   a. Crises are rare
   b. In novel, high-stakes moments AI is least reliable and human judgment most needed — exactly when debt comes due
   c. Regulations require it
   d. Boards prefer it

**12. "Every convenience is a loan against a skill" implies the leadership discipline of…**

   a. Simply refusing all convenience
   b. Choosing which skills to keep
   c. Banning AI use for all juniors
   d. Regular weekly memory tests for all

### Module 15 — The AI Paradox: Adoption Without Transformation

**Outcome:** Diagnose why high AI adoption so often yields low transformation — and design change strategies that beat organisational antibodies.

**Leadership lens:** "Everyone uses Copilot" and "nothing has changed" are both true in most enterprises. AI theater, structural mismatch, and middle-management antibodies are diagnosable and treatable — if a leader is willing to name them.

**Apply-at-work mission — Run an adoption audit:** Conduct an honest AI Adoption Audit of your team or organisation: where is adoption high but transformation low? Name the top 3 organisational antibodies and one countermeasure for each.

**Reflection:** Where am I personally performing AI theater — visible adoption without changed outcomes? What would real transformation of my own role look like?

#### Resources

- [McKinsey QuantumBlack — why AI projects fail](https://www.mckinsey.com/capabilities/quantumblack/our-insights) — The transformation-gap evidence: high adoption, low P&L impact — and the diagnosed causes.
  - Read the high-adoption, low-impact evidence and its diagnosed causes
  - Diagnose the transformation gap in your own organisation
  - Separate adoption metrics from P&L impact
- [HBR — change management (collection)](https://hbr.org/topic/change-management) — Classic change frameworks; read them asking "what breaks when the change is continuous?"
  - Ask what breaks when change is continuous, not one-off
  - Read classic frameworks against a never-stable environment
  - Adapt change management for permanent iteration
- [a16z — State of AI (industry analysis)](https://a16z.com/state-of-ai/) — The investor view of adoption vs transformation: where real workflow change is happening.
  - Locate where real workflow change is happening versus surface adoption
  - Take the investor read on transformation
  - Tell genuine change from tool-adoption theatre

#### Project — AI Adoption Audit + Change Strategy

Conduct an honest AI Adoption Audit of your team or organisation: where is adoption high (tools bought, accounts active) but transformation low (workflows, structures, and outcomes unchanged)? Identify the top 3 "organisational antibodies" resisting deeper agentic adoption in your context — and design a change strategy addressing both technical capability and organisational dynamics, with one countermeasure per antibody.

**Deliverable:** portfolio/w15-adoption-audit.md — the audit findings, 3 antibodies with evidence, and the change strategy.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Audit rigour | 30% | Adoption vs transformation distinguished with observable evidence (usage stats vs changed processes/KPIs), not impressions. |
| Antibody diagnosis | 25% | The 3 antibodies are specific to your organisation (e.g. utilisation-based billing, review bottlenecks, role-protection in layer X) — with the incentive behind each named. |
| Change strategy realism | 30% | Countermeasures address incentives and structure, not just communication; each has an owner, a first step, and a way to tell it's working. |
| Self-inclusion | 15% | Your own AI theater identified — where your visible adoption exceeds your changed outcomes. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. The "AI paradox" of adoption without transformation describes…**

   a. Nobody in the whole company using AI
   b. High usage, unchanged outcomes
   c. AI quietly replacing every single person
   d. A totally failed procurement process

**2. "Organisational antibodies" are…**

   a. A kind of security software
   b. Things that block change
   c. Certain difficult employees
   d. The various compliance rules

**3. A classic structural mismatch blocking transformation is…**

   a. The offices are simply far too small for all of it
   b. Human-speed structures around machine-speed work
   c. Too few licences
   d. Wrong cloud region

**4. Middle-management resistance to agentic transformation is usually driven by…**

   a. Widespread technophobia among the managers
   b. Rational role protection by managers
   c. Certain age demographics being at play
   d. Union rules

**5. "AI theater" is…**

   a. Corporate training videos
   b. Visible AI activity only
   c. AI used in the entertainment industry
   d. Long vendor presentations

**6. The most reliable test distinguishing transformation from adoption is…**

   a. Survey sentiment
   b. Whether workflows, decision rights, or KPIs have structurally changed — could the old process still run unchanged?
   c. Licence utilisation
   d. Executive quotes

**7. Why do communication-only change programmes fail against antibodies?**

   a. Emails go unread
   b. Antibodies live in incentives and structures; persuasion doesn't change what people are paid and promoted for
   c. Wrong channels
   d. Too few townhalls

**8. A countermeasure for "utilisation-billed teams resist efficiency gains" is…**

   a. More training
   b. Change the incentive model: price outcomes
   c. Simply mandate all AI usage across the board
   d. Hire consultants

**9. Real capability building (vs theater) looks like…**

   a. A regular internal company AI newsletter
   b. Redesigned workflows in production
   c. A big internal prompt-writing competition
   d. An innovation lab tour

**10. Including YOURSELF in the adoption audit matters because…**

   a. Humility is fashionable
   b. Leaders' own theater licenses everyone else's; credibility to demand change requires having made it personally
   c. Auditors require it
   d. It softens findings

**11. The correct sequencing for beating antibodies is usually…**

   a. A big company-wide mandate right on the first day
   b. Prove value in one real workflow, then expand
   c. Wait for competitors
   d. Reorg first

**12. High adoption with low transformation is DANGEROUS (not merely disappointing) because…**

   a. The software licences all expire
   b. It inoculates the organisation
   c. The vendors all get rich from it
   d. IT gets blamed

---

## Phase 5 — Responsible Scale & Capstone (weeks 16–18)

### Module 16 — Responsible AI and Agentic Scale

**Outcome:** Build governance for systems that act: guardrails + monitoring over approval gates, with trust, recourse, and rollback at scale.

**Leadership lens:** A hallucination in a chatbot is an embarrassment; a hallucination in an agent with system access is an incident. Governance design — autonomy tiers, monitoring, escalation, rollback — is the steward's core craft.

**Apply-at-work mission — Draft a real risk register:** Take one live or proposed agentic workflow in your organisation and produce a Risk Register + governance one-pager: failure modes, autonomy tiers, monitoring signals, escalation path, rollback plan.

**Reflection:** What's the riskiest thing my organisation currently lets AI do with no monitoring? Why has nobody asked — and what does it cost me to be the one who asks?

#### Resources

- [NIST — AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework) — The reference governance vocabulary: map, measure, manage, govern. Skim the core functions.
  - Use map / measure / manage / govern vocabulary in governance design
  - Skim the core functions well enough to apply them
  - Frame governance in terms auditors accept
- [Agentic AI safety & governance research (arXiv)](https://arxiv.org/search/?query=agentic+ai+safety+governance) — Recent work on the failure modes unique to systems that act, not just answer.
  - Name the failure modes unique to systems that act, not just answer
  - Anticipate agent-specific risks in your design
  - Read current safety work for what it changes
- [LangSmith — LLM observability docs](https://docs.smith.langchain.com/) — What production monitoring of AI systems actually looks like: traces, evals, feedback loops.
  - Know what production monitoring looks like: traces, evals, feedback loops
  - Design monitoring over approval gates for systems that act
  - Specify the observability your governance needs
- [Anthropic — guardrails & safety research](https://www.anthropic.com/research) — How guardrails are built at the model and system level — the technical backdrop to your governance design.
  - Understand how guardrails are built at model and system level
  - Ground your governance design in real guardrail mechanics
  - Set what a guardrail can and cannot guarantee

#### Project — Governance Framework + Risk Register (Capstone Milestone 5)

Design a governance framework for agentic workflows in your organisation: autonomy tiers (what agents may do freely / with approval / never), monitoring signals, escalation paths, and rollback procedures. Build a Risk Register for one real or proposed agentic system covering hallucination, action, and compliance risks — each with likelihood, impact, and a named control. If you can, add a simple monitoring/logging view for one of your own workflows. 🎯 This completes Capstone Milestone 5: governance you could table at a risk committee.

**Deliverable:** portfolio/w16-governance.md — the framework one-pager, the risk register, and (optional) monitoring screenshots.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Autonomy tier design | 25% | Three-plus tiers with concrete examples per tier from your context; tier boundaries justified by reversibility and stakes. |
| Risk register quality | 30% | 8+ real risks across hallucination/action/compliance, each with likelihood, impact, owner, and a control that would actually work. |
| Monitoring & escalation | 25% | Named signals (error rate, cost, drift, complaint rate) with thresholds, plus who gets paged and what they do. |
| Rollback realism | 20% | A stop-and-recover procedure that works at 2am without the builder: kill switch, fallback process, communication plan. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. Agentic systems carry categorically higher risk than chatbots because…**

   a. They cost more
   b. Hallucination plus action: errors become side effects
   c. They tend to run more or less continuously all the time
   d. They use more tokens

**2. Autonomy tiers structure governance by…**

   a. Ranking agents by IQ
   b. Defining what actions run freely, which need approval, and which are forbidden — by stakes and reversibility
   c. Pricing model access
   d. Separating dev and prod

**3. The shift "from approval gates to guardrails + monitoring" happens because…**

   a. Approvals are expensive
   b. At scale, per-action human approval collapses; continuous constraints + observability keep control without the bottleneck
   c. Regulators changed rules
   d. Guardrails sound better

**4. A good monitoring SIGNAL for an agentic workflow is…**

   a. The overall server uptime figure alone
   b. Behavioural rates, not just uptime
   c. The total number of active users only
   d. Model version

**5. A risk register entry is complete when it has…**

   a. A suitably scary-sounding name
   b. An owner and a control
   c. A neat colour-coded label
   d. A full insurance quote for it

**6. "Recourse" as a trust property means…**

   a. A particular form of legal insurance
   b. Affected parties can reach a human
   c. Standard company refund policies in place
   d. Apology templates

**7. The rollback plan must work "at 2am without the builder" because…**

   a. Builders sleep late
   b. Autonomous systems fail on their own schedule
   c. The auditors always run all their tests at night
   d. It sounds rigorous

**8. Portfolio governance for agentic initiatives means…**

   a. One mega-project
   b. Managing the whole set with shared standards
   c. A single steering committee with its own logo
   d. Annual reviews

**9. Transparency about agent involvement (disclosure) matters because…**

   a. The marketing team generally prefers it
   b. Discovered concealment destroys trust
   c. It is always strictly legally required
   d. Users read policies

**10. Which failure is an ACTION risk (vs hallucination risk)?**

   a. Inventing a statistic in a summary
   b. A wrongly sent refund
   c. Citing the wrong page number
   d. Overly verbose answers

**11. An agent's behaviour slowly drifts as usage patterns change. The governance control is…**

   a. Just simply hoping for the best
   b. Baseline metrics and alerts
   c. Restarting the whole thing weekly
   d. Using much bigger prompts throughout

**12. Presenting governance as an ENABLER (not a brake) is credible when…**

   a. It has a nice logo
   b. Clear tiers and pre-approved guardrails let teams ship low-risk automation FASTER, with escalation only where stakes demand
   c. It is optional
   d. Fines are mentioned

### Module 17 — Continuous Transformation: Leading When AI Never Stops Evolving

**Outcome:** Build the operating rhythm — funding, teams, iteration cadence — for an organisation where the capability curve never flattens.

**Leadership lens:** This transformation has no "done". AI factories, agent-to-agent ecosystems, persistent iteration teams — leaders must normalise permanent beta without exhausting their people. That's an operating-model design problem.

**Apply-at-work mission — Pitch an iteration team:** Draft and pitch (to a real stakeholder, even informally) a lightweight funding + governance model for a persistent AI iteration capability in your organisation: who, budget envelope, cadence, kill criteria.

**Reflection:** What's my personal operating rhythm for staying current without drowning? What did I stop doing to make room for it?

#### Resources

- [Mustafa Suleyman — The Coming Wave (key ideas)](https://www.youtube.com/results?search_query=mustafa+suleyman+the+coming+wave+summary) — Containment, proliferation, and why this wave doesn't crest — the strategic backdrop for permanent iteration.
  - Understand containment and proliferation as the strategic backdrop
  - See why this wave does not crest
  - Plan for permanent iteration, not a finish line
- [HBR — organisational learning (collection)](https://hbr.org/topic/organizational-learning) — Senge and successors: learning organisations, updated for an environment that never stabilises.
  - Update learning-organisation ideas for a never-stable environment
  - Design an operating rhythm that keeps learning
  - Build the funding and cadence for continuous change
- [Future of agentic AI — research & predictions (arXiv)](https://arxiv.org/search/?query=future+agentic+ai+2026) — Scan the frontier: agent-to-agent protocols, AI factories, ecosystem plays — inputs for your scenarios.
  - Scan agent-to-agent protocols, AI factories and ecosystem plays
  - Gather inputs for your own scenario planning
  - Read the frontier for strategy, not novelty

#### Project — Continuous Transformation Playbook

Write the "Continuous Transformation Playbook" for your team or organisation: operating rhythm for ongoing AI iteration, a lightweight funding + governance model for a persistent iteration capability (who, budget envelope, cadence, kill criteria), and 3 scenarios for how agentic AI reshapes your industry/function in 2028–2030 — each with a "we would start doing X now" implication. Pitch the iteration-team model to a real stakeholder, even informally, and record the reaction.

**Deliverable:** portfolio/w17-transformation-playbook.md — the playbook, funding model, 3 scenarios, and the pitch outcome.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Operating rhythm | 25% | A concrete cadence (weekly/monthly/quarterly loops) naming what gets evaluated, adopted, retired — sustainable, not heroic. |
| Funding & governance model | 30% | A real proposal: team shape, budget envelope, decision rights, and explicit kill criteria for experiments. |
| Scenario quality | 25% | Three genuinely different 2028–2030 futures with present-tense implications — each names something to START now. |
| Real pitch | 20% | You pitched a real person; their objections and your revisions documented. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. AI transformation differs from ERP/cloud transformations because…**

   a. It is generally cheaper
   b. It has no end state
   c. It needs no real training
   d. The vendors are all nicer

**2. A persistent "AI iteration team" exists to…**

   a. Write out all of the prompts for everyone else here too
   b. Continuously evaluate, upgrade, and retire workflows
   c. Run the helpdesk
   d. Negotiate licences

**3. Kill criteria for AI experiments matter because…**

   a. The legal team strictly demands them
   b. Zombie pilots drain the budget
   c. They really do motivate the teams
   d. The boards all rather like the word

**4. "Normalising iteration" for live agentic systems means…**

   a. Holding weekly all-hands meetings
   b. Treating updates as routine
   c. Freezing all the systems yearly
   d. Blaming users less

**5. Funding persistent iteration as OPEX/product (vs one-off project CAPEX) matters because…**

   a. The accounting department simply always prefers it
   b. Project money ends and the capability decays
   c. It is cheaper
   d. CFOs never ask about OPEX

**6. "AI factories" as a future trend refers to…**

   a. Chip fabrication plants
   b. Organisations industrialising the production of AI-powered workflows — pipelines that turn processes into agentic systems repeatably
   c. Robot warehouses
   d. Prompt marketplaces

**7. Agent-to-agent (A2A) protocols would strategically matter because…**

   a. They cut API costs
   b. Your agents will transact with suppliers' and customers' agents — interfaces, trust, and standards become competitive terrain
   c. They are faster
   d. They simplify logging

**8. Scenario planning beats point predictions for 2028–2030 because…**

   a. Predictions are boring
   b. Under deep uncertainty, preparing for several divergent futures builds capabilities that pay off across most of them
   c. Scenarios need no data
   d. Consultants sell them

**9. The biggest HUMAN risk of permanent transformation is…**

   a. General wage and salary inflation
   b. Change fatigue and burnout
   c. Simply far too many meetings
   d. An overall surplus of skills

**10. A "learning organisation" in the agentic era is distinguished by…**

   a. A very big annual training budget
   b. Short feedback loops, in weeks
   c. Regular industry conference attendance
   d. An LMS platform

**11. Your personal operating rhythm for staying current should optimise for…**

   a. Reading through absolutely everything
   b. Filtered, recurring, bounded signal
   c. Constantly following the influencers
   d. Daily tool-switching

**12. Pitching the iteration model to a real stakeholder (this week's mission) matters because…**

   a. Practice makes perfect
   b. Real objections reveal actual constraints
   c. The stakeholders genuinely enjoy the pitches
   d. It completes the rubric

### Module 18 — Capstone: Build, Govern, and Tell the Story

**Outcome:** Synthesise everything: identify, design, build, and govern a real agentic AI solution — then present it like a leader.

**Leadership lens:** The capstone is your proof-of-leadership: a real problem, a working (or credibly prototyped) agentic solution, a governance posture, measured impact, and a story that lands with executives and engineers alike.

**Apply-at-work mission — Ship the capstone:** Complete your capstone in the Capstone Tracker: real problem, built solution or detailed prototype, governance + risk register, impact measurement, published portfolio, and a 10–15 min recorded walkthrough.

**Reflection:** Final entry: read your Week 1 reflection, then write to your past self. What did you most misunderstand about leading with AI 18 weeks ago?

#### Resources

- [Building a strong AI portfolio (curated guides)](https://www.youtube.com/results?search_query=building+ai+portfolio+github) — How to present technical-adjacent work credibly: structure, evidence, storytelling.
  - Present technical-adjacent work with structure and evidence
  - Tell the story of a system credibly
  - Structure a portfolio that lands with non-technical leaders
- [HBR — communicating technical work to stakeholders](https://hbr.org/topic/communication) — The presentation half of the capstone: translating systems into business narrative.
  - Translate a system into a business narrative
  - Lead the capstone presentation with impact, not mechanics
  - Make your work legible to stakeholders
- [All tools from previous weeks](https://docs.n8n.io/) — n8n, CrewAI, Dify, Claude, Cursor — your capstone stack is everything you've already used.
  - Assemble your capstone stack from everything you have already used
  - Combine n8n, CrewAI, Dify, Claude and Cursor into one solution
  - Reuse rather than restart for the capstone
- [GitHub / Notion — portfolio hosting](https://github.com/) — Where the portfolio lives: public repo or polished Notion page — linkable, professional, current.
  - Host a linkable, professional, current portfolio
  - Choose between a public repo and a polished Notion page
  - Make your work findable and shareable

#### Project — Capstone: Build, Govern, and Tell the Story (Capstone Milestone 6)

Identify a meaningful real problem in your work or community. Design and build a complete agentic AI solution (or detailed prototype) addressing it — including governance, monitoring, and scaling considerations from your weeks 6, 8, and 16 artifacts. Create a professional portfolio (site or Notion) showcasing the capstone, 3–4 other programme projects, and your leadership philosophy. Record a 10–15 minute video walkthrough. 🎯 This completes Capstone Milestone 6 — and the programme.

**Deliverable:** portfolio/w18-capstone.md + published portfolio link + video — the complete story: problem, solution, governance, impact, journey.

**Assessment rubric**

| Criterion | Weight | What good looks like |
| --- | ---: | --- |
| Problem significance & fit | 20% | A real problem someone actually has, matched honestly to agentic capability — not a solution seeking a problem. |
| Solution completeness | 30% | Working system or credible detailed prototype; components, autonomy tiers, and human gates all deliberate and explained. |
| Governance & impact | 25% | Risk register, monitoring, and scaling plan attached; impact measured or honestly estimated with a method. |
| Portfolio & narrative | 25% | Published portfolio a stranger could assess; the video tells a leadership story — problem, judgment calls, results — not a feature tour. |

#### Knowledge check (12 questions)

**Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.**

**1. The strongest capstone problem selection criterion is…**

   a. Sheer technical impressiveness
   b. A real, well-fitting pain
   c. Trending new technology use
   d. Maximum possible agent count

**2. A capstone with brilliant engineering but NO governance section signals…**

   a. Efficient prioritisation
   b. Incomplete leadership: capability without accountability is exactly what this programme argues against
   c. Trust in the model
   d. Agile thinking

**3. Measuring capstone impact honestly, when full data isn't available, means…**

   a. Simply skipping impact claims
   b. A method and honest estimate
   c. Using the biggest defensible number
   d. Quoting from various analyst reports

**4. Your portfolio's primary audience design should target…**

   a. Various other programme graduates just like you
   b. A sceptical decision-maker assessing you fast
   c. Search engines
   d. AI researchers

**5. The video walkthrough should be structured as…**

   a. A feature-by-feature screen tour
   b. A leadership narrative
   c. A full tutorial for rebuilding it
   d. A slide deck reading

**6. "Leadership story" framing of your 18 weeks means…**

   a. Simply listing all the completed modules
   b. Articulating how your thinking changed
   c. A plain chronological diary of the whole thing
   d. Certificates displayed

**7. Including FAILURES in the portfolio (what didn't work) is…**

   a. Career suicide
   b. A credibility multiplier: honest failure analysis
   c. It is really just some entirely optional padding for it
   d. Only for engineers

**8. The multi-agent SDR system, ops command center, and governance dashboard capstone directions share…**

   a. They all happen to use the very same tools
   b. End-to-end scope including oversight
   c. A strong shared sales-department focus
   d. Low difficulty

**9. Presenting to executives vs engineers differs in that executives primarily need…**

   a. A great many more acronyms used
   b. Outcomes, risks, decisions
   c. Architecture diagrams first
   d. Much longer meeting sessions

**10. The "letter to your Week-1 self" reflection exercise exists because…**

   a. Sentimentality
   b. Articulating your own change consolidates it — and reveals the misconceptions you'll now recognise in others
   c. It fills the last week
   d. Tradition

**11. After the programme, the highest-value habit to keep is…**

   a. Weekly quiz-taking
   b. The build-reflect-apply loop: hands-on testing of new capabilities, journaled reflection, and workplace application
   c. Reading more newsletters
   d. Collecting certificates

**12. Sharing your capstone publicly ("build in public") primarily buys you…**

   a. Likes
   b. Compounding credibility and a network
   c. A definite risk of plagiarism by others
   d. Nothing measurable

## Toolkits

### Agentic Workflow Canvas

**Unlocks in module 6.**

Map any process for agents: inputs, decisions, tools, exceptions, and the human gates that stay.

```markdown
# Agentic Workflow Canvas

**Process name:**   **Owner:**   **Date:**

## 1. Outcome
What does "done, and done well" look like? Who consumes the output?

## 2. Trigger & inputs
What starts the process? List every input, its source, and its format (including the messy ones).

## 3. Steps & decisions
| # | Step | Actor (human/agent) | Decision criteria (if a decision) | Tools/data used |
|---|------|--------------------|-----------------------------------|-----------------|
| 1 | | | | |

## 4. Exception paths
For every step: what can go wrong, how is it detected, and where does it route (retry / fallback / escalate)?

| Step | Failure mode | Detection | Route |
|------|-------------|-----------|-------|

## 5. Human judgment gates (minimum 5 candidates, keep the real ones)
| Gate | Why human? (stakes / ambiguity / irreversibility) | What the human sees to decide |
|------|---------------------------------------------------|-------------------------------|

## 6. Observability
What is logged at each step? How would you reconstruct a run a week later?

## 7. Kill switch
How do you stop this workflow fast, and what is the manual fallback while it's stopped?
```

### Pilot-to-Scale Maturity Assessment

**Unlocks in module 8.**

Grade any AI pilot on the 4-level model and plan the jump to the next level.

```markdown
# Maturity Assessment — Experiment → Pilot → Operational → Scaled

**Workflow/pilot:**   **Assessed by:**   **Date:**

## Level definitions
1. **Experiment** — works on curated examples, builder-operated, no owner
2. **Pilot** — real users, real inputs, builder still on call, success criteria defined
3. **Operational** — named owner, error handling, integrated into real systems, support process exists
4. **Scaled** — multiple teams/contexts, portfolio governance, funded as a product

## Assessment
| Dimension | Evidence today | Level (1–4) |
|-----------|---------------|-------------|
| Ownership (who is accountable when it breaks?) | | |
| Input reality (curated vs whatever arrives) | | |
| Error handling & exceptions | | |
| Integration (real systems vs copy-paste) | | |
| Monitoring & metrics | | |
| Funding model (project vs product) | | |

**Overall current level:**   **Target level (6 months):**

## Blockers to next level (be specific)
1.
2.
3.

## Redesigned SOP (attach)
What does the agent do, what do humans do, how do exceptions route?

## Human+agent KPIs (efficiency AND quality)
| KPI | Type | Measurement method | Target |
|-----|------|--------------------|--------|
```

### Governance Playbook & Risk Register

**Unlocks in module 16.**

Autonomy tiers, monitoring signals, escalation, rollback — plus the risk register for any agentic system.

```markdown
# Agentic Governance Playbook & Risk Register

**System:**   **Owner:**   **Steward:**   **Date:**

## Autonomy tiers
| Tier | Definition | Examples for this system |
|------|-----------|--------------------------|
| Free | Agent acts, logs, no approval | |
| Gated | Agent prepares, human releases | |
| Forbidden | Agent may never do this | |

## Monitoring signals
| Signal | Threshold | Who is alerted | First response |
|--------|-----------|----------------|----------------|
| Error / correction rate | | | |
| Escalation rate | | | |
| Cost per run / day | | | |
| Behaviour drift vs baseline | | | |

## Escalation path
Who gets paged, in what order, with what authority?

## Rollback procedure (must work at 2am without the builder)
1. Kill switch location + who may pull it:
2. Manual fallback process while stopped:
3. Communication plan (users, stakeholders):

## Risk register
| # | Risk | Type (hallucination / action / compliance) | Likelihood | Impact | Owner | Control |
|---|------|--------------------------------------------|-----------|--------|-------|---------|
| 1 | | | | | | |
| 2 | | | | | | |
| 3 | | | | | | |

## Recourse
How does an affected person challenge an outcome, and which human can change it?
```

### Cognitive Capital Audit

**Unlocks in module 14.**

Classify two weeks of your AI usage: thinking deepened (capital) or replaced (debt)?

```markdown
# Cognitive Capital Audit

**Period reviewed:**   **Date:**

## Usage log
| Use of AI (significant instances) | What I did first, myself | Capital or Debt? | Evidence |
|-----------------------------------|--------------------------|------------------|----------|
| | | | |

**Capital** = my understanding/skill is stronger after this pattern of use.
**Debt** = I can no longer comfortably do (or judge) this without the tool.

## Debt areas (top 3)
| Skill degrading | Evidence | Do I want it back? | Repayment plan (practice) |
|-----------------|----------|--------------------|-----------------------------|

## My Thinking Amplification Protocol
Trigger (which decisions/documents):
1. Draft my own position first (timebox:   min)
2. Ask AI to attack it: weakest assumptions, missing perspectives, disconfirming evidence
3. Revise and record what changed
Practice commitment:   days/week for   weeks

## Team norms I will model
- Think-then-ask on:
- "The AI said so" is never a complete justification for:
```

### Orchestrator–Architect–Steward Role Map

**Unlocks in module 13.**

Map your team against the three roles that scale; find the gap; draft the missing job.

```markdown
# Role Mapping — Orchestrator / Architect / Steward

**Team:**   **Mapped by:**   **Date:**

## Role definitions
- **Orchestrator** — aligns humans + agents across silos; owns handoffs, context, and cadence
- **Architect** — designs workflows, decision logic, and system structure
- **Steward** — owns trust, safety, accountability, and governance of human+agent systems

## Current coverage
| Person / function | Orchestrator | Architect | Steward | Evidence |
|-------------------|--------------|-----------|---------|----------|
| | | | | |

## Failure symptoms observed (map to missing role)
| Symptom | Points to gap in |
|---------|------------------|
| Governance incidents despite fast shipping | Steward |
| Fragile ad-hoc undocumented workflows | Architect |
| Silos, duplicated agent work, broken handoffs | Orchestrator |
| (your observations) | |

## My gravitational role
Under pressure I default to:   Evidence:
Development plan for my weakest role:

## Draft job description for the missing role
Title:
Mission:
Responsibilities (5):
Success at 6 months / 12 months:
```

### Continuous Transformation Playbook

**Unlocks in module 17.**

Operating rhythm, funding model, and kill criteria for a persistent AI iteration capability.

```markdown
# Continuous Transformation Playbook

**Scope (team/org):**   **Author:**   **Date:**

## Operating rhythm
| Cadence | Activity | Output |
|---------|----------|--------|
| Weekly | Frontier scan + one hands-on test | Tested-capability note |
| Monthly | Evaluate 1–2 capabilities against live workflows | Adopt / watch / reject decision |
| Quarterly | Portfolio review: retire, scale, re-govern | Updated workflow portfolio |

## Iteration capability
Team shape (roles, % time):
Budget envelope:
Decision rights (what they may change without approval):

## Kill criteria (pre-committed)
An experiment stops when:   (e.g. no measurable value after N cycles, cost/quality regression, unowned risk)

## Change-fatigue guards
Stable anchors that do NOT change (purpose, quality bar, core rituals):

## Scenarios 2028–2030 (three futures)
| Scenario | What the industry looks like | What we start doing NOW |
|----------|------------------------------|--------------------------|
| | | |
```

### Apply-at-Work Mission Log

**Unlocks in module 1.**

The running record of every weekly mission: what you did at work, what happened, what you learned.

```markdown
# Apply-at-Work Mission Log

| Week | Mission | What I actually did | Outcome / reaction | What I'd do differently |
|------|---------|---------------------|--------------------|-------------------------|
| 1 | Hire your first AI teammate | | | |
| 2 | Brief your team on the inflection | | | |
| 3 | Automate one monitoring task | | | |
| 4 | Ship a tool your team uses | | | |
| 5 | Connect a prototype to a real system | | | |
| 6 | Canvas a real workflow with a colleague | | | |
| 7 | Stand up a multi-agent system | | | |
| 8 | Maturity-assess a real pilot | | | |
| 9 | Build a lead/ops research agent | | | |
| 10 | Repurpose content with brand guardrails | | | |
| 11 | Design a review-gated HR/Finance workflow | | | |
| 12 | Run a WHY–WHAT–HOW conversation | | | |
| 13 | Role-map your team | | | |
| 14 | Audit your AI usage | | | |
| 15 | Run an adoption audit | | | |
| 16 | Draft a real risk register | | | |
| 17 | Pitch an iteration team | | | |
| 18 | Ship the capstone | | | |
```

## Capstone

### Enterprise · Multi-agent SDR system

Research + qualification + personalised outreach preparation, with compliance gates and a human release step.

### Enterprise · Intelligent operations command center

Monitors multiple data sources, correlates signals, and triggers coordinated (gated) responses.

### Enterprise · Customer success intervention system

Detects churn-risk signals and orchestrates intervention workflows with human owners.

### Leadership & Governance · Agentic governance framework + monitoring dashboard

A real organisation's autonomy tiers, risk register, and a live monitoring view over its agent workflows.

### Leadership & Governance · Role-based agent team design for a department

Orchestrator + Architect + Steward blueprint applied to one department, with workflows and decision rights.

### Leadership & Governance · Cognitive capital development programme

A team-level programme using AI as thinking partner (not answer engine): protocols, norms, and measurement.

### Your own · Your own real problem

The best capstone: a meaningful problem from your work or community that an agentic approach genuinely fits.

### Milestones

- **Module 3 — First working agentic workflow**
- **Module 6 — Agentic workflow canvas for a real process**
- **Module 8 — Pilot-to-scale plan**
- **Module 13 — Role map + agentic team blueprint**
- **Module 16 — Governance framework + risk register**
- **Module 18 — Capstone shipped + portfolio published**

### Portfolio checklist

- Capstone case study: problem → solution → governance → impact
- 3–4 supporting projects from the programme (weeks 3, 6, 8, 13, 16)
- Leadership philosophy / manifesto (Week 12, revised)
- Leadership story: how my thinking changed across 18 weeks
- Video walkthrough (10–15 min) linked
- LinkedIn-ready summary post drafted
