Agentic AI for Leadership

Executive Certificate — 18 weeks, self-paced, modelled on the NTU Singapore Executive Certificate in Agentic AI Leadership

This file is generated from the course data by scripts/build-notes.mjs. Edit the course data, not this file.


Phase 1 — Foundations: From Generative to Agentic (weeks 1–3)

Module 1 — From Search to Digital Teammates

Outcome: Experience the shift from AI-as-search-tool to AI-as-collaborative-teammate — and feel how your own workflow changes.

Leadership lens: Your calendar is full of work a persistent teammate could carry: briefing packs, meeting prep, first drafts. The leaders who win aren't the best prompters — they're the best delegators to a new kind of colleague.

Apply-at-work mission — Hire your first AI teammate: Stand up a persistent AI teammate (Claude Project / Custom GPT / NotebookLM) for ONE recurring task in your actual job — give it a role, goals, and your working style. Use it every working day this week.

Reflection: How does my thinking change when I treat AI as a teammate instead of a tool? What did I delegate this week that I would never have delegated a month ago?

Resources

Project — Your First AI Teammate

Create a persistent AI teammate using Claude Projects or a Custom GPT: give it a clear role, goals, and knowledge of your working style. Separately, load 10–15 documents about your industry or function into NotebookLM and hold three deep conversations with your "knowledge teammate". Then write a 500-word reflection: "How my thinking changes when I treat AI as a teammate instead of a tool."

Deliverable: portfolio/w01-ai-teammate.md — teammate setup (role prompt), 3 conversation takeaways, and the 500-word reflection.

Assessment rubric

Criterion Weight What good looks like
Teammate design 25% The role prompt defines role, goals, tone, and working context specifically enough that a colleague could tell whose teammate it is; not a generic "helpful assistant".
Real daily use 25% Evidence of use on 5+ real work items across the week (drafts, prep, analysis), with notes on what worked and what disappointed.
Knowledge-teammate experiment 20% 10–15 genuinely relevant documents in NotebookLM; three conversations that surface at least one insight you did not already have.
Reflection quality 30% The 500 words name a concrete change in how you think or delegate — not "AI is amazing" but a specific before/after in your own behaviour.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. What most fundamentally distinguishes an AI "teammate" from an AI "tool"?

  1. Bigger model size
  2. Persistence, memory, context, and initiative
  3. Considerably faster average response times overall
  4. Access to the internet

2. The evolution this module traces is best described as…

  1. Search engines → databases → spreadsheets
  2. Search → copilots → autonomous agents
  3. Chatbots → progressively bigger chatbots
  4. On-prem → to cloud → to the edge over time

3. Why does giving an AI a standing role and goals ("persistent context") matter for leaders?

  1. It reduces API costs
  2. It turns each interaction into true delegation
  3. It makes the model considerably more accurate at maths
  4. It is required by vendors

4. A leader uploads 15 strategy documents to NotebookLM and interrogates them. This practice primarily builds…

  1. A full production-grade RAG system
  2. A personal knowledge teammate
  3. A properly fine-tuned custom model
  4. An automation pipeline of some kind

5. Which behaviour signals someone still treats AI as a tool, not a teammate?

  1. Reviewing all AI output before using it
  2. Giving it a clear role and feedback over time
  3. Re-explaining context each time
  4. Delegating your meeting prep to it

6. The "experience layer" shift (voice, multimodal, always-available AI) matters to leaders mainly because…

  1. It looks modern in demos
  2. It changes when and where thinking work happens, not just how fast
  3. It reduces licence costs
  4. It removes the need for meetings

7. What is the biggest realistic risk of week 1's "AI teammate" practice?

  1. The AI simply refuses to do any of the work
  2. Accepting fluent output without judgment
  3. Quietly running out of all your available tokens
  4. Colleagues noticing

8. Why do many professionals report AI "doesn't help much" after trying it once?

  1. The models are weak
  2. They issued one-shot queries with no context
  3. AI only helps engineers
  4. Their privacy settings quietly block the quality

9. A well-designed teammate role prompt should include…

  1. Only the single small task of the day itself
  2. Role, goals, audience, tone, and style
  3. As little as you possibly can, to avoid bias
  4. Technical model parameters

10. The reflection exercise ("how my thinking changes") exists because…

  1. Writing fills the week
  2. Leadership development happens through reflection on practice, not consumption of content
  3. It produces content for LinkedIn
  4. It tests writing skills

11. Which task is the WEAKEST first delegation to a new AI teammate?

  1. Drafting up a meeting agenda from your notes
  2. Summarising a rather long and detailed report
  3. A final, unreviewed client message
  4. Preparing interview questions

12. The healthiest mental model for current AI teammates is…

  1. An infallible oracle
  2. A fast, tireless junior colleague with confidence issues in reverse — always confident, sometimes wrong
  3. A search engine with better grammar
  4. A threat to be minimised

Module 2 — Why Now: The AI Inflection Point

Outcome: Explain the technical and economic forces behind the current inflection — and why agentic AI is not another RPA wave.

Leadership lens: Boards ask "why now, why us, why this much?" This week gives you the cost-collapse, capability-explosion, and time-to-X arguments to answer in their language: valuation, revenue per employee, competitive moats.

Apply-at-work mission — Brief your team on the inflection: Run a 15-minute briefing for your team or manager: one chart, one industry example, one implication for your business in the next 12 months. Note what convinced them and what met resistance.

Reflection: Which part of my industry's value chain is most exposed to the time-to-X collapse — and am I positioned as a spectator or an architect of that change?

Resources

Project — Industry Disruption Analysis

Choose one industry you know deeply. Write a 2-page analysis: "How the AI inflection point is already disrupting (or will disrupt) this industry in the next 3 years" — covering cost collapse, capability explosion, and the time-to-X collapse. Document 3 companies using agentic approaches today, and build a simple AI Inflection Timeline (2022–2027) of the capability jumps that matter to this industry.

Deliverable: portfolio/w02-inflection-analysis.md — the 2-page analysis, 3 company snapshots, and the timeline.

Assessment rubric

Criterion Weight What good looks like
Economic argument 30% Uses real numbers (cost per token/task trends, adoption data) rather than vibes; distinguishes efficiency gains from business-model change.
Agentic vs previous waves 20% Explains concretely why this differs from RPA/traditional ML for THIS industry — what becomes possible, not just cheaper.
Company evidence 25% Three real companies with sourced descriptions of their agentic approach and observable results; no press-release-only examples.
Timeline & judgment 25% Timeline picks capability jumps relevant to the industry, and the analysis takes a position — where value moves, who is exposed, what a leader should do in the next 2 quarters.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. The two simultaneous forces defining the current inflection point are…

  1. Rising costs and shrinking capability
  2. Falling cost, rising capability
  3. More regulation and much less compute
  4. Fewer models but many more vendors

2. "Time-to-X collapse" refers to…

  1. Servers simply responding faster
  2. The shrinking idea-to-output gap
  3. Noticeably shorter overall working hours
  4. Faster model training runs overall

3. Agentic AI differs from RPA fundamentally because…

  1. It is cheaper to license
  2. RPA replays fixed clicks; agents adapt
  3. The agents only ever work up in the cloud
  4. RPA is newer

4. "AI-native economics" most refers to…

  1. Buying GPUs in bulk
  2. Business models built assuming near-zero marginal cost of cognition — e.g. revenue per employee outliers
  3. Charging more for AI features
  4. Paying engineers in equity

5. Why is "waiting for the next budget cycle" newly dangerous at an inflection point?

  1. Budgets are illegal to delay
  2. Capability compounds while rivals accelerate
  3. The vendors always raise their prices annually
  4. Auditors disapprove

6. The most defensible way to size AI's impact on YOUR industry is…

  1. Quoting the biggest consulting number
  2. Mapping specific value-chain steps to capabilities
  3. Simply waiting for a peer to publish their results first
  4. Running an employee survey

7. A "capability jump" worth putting on your timeline is one that…

  1. Made headlines
  2. Crossed a reliability/cost threshold that unlocks a real workflow in your industry
  3. Won a benchmark
  4. Was demoed at a conference

8. Foundation models differ from traditional ML projects in that they…

  1. They need more labelled data per task
  2. General-purpose, adapted per task
  3. They only ever work for spoken language
  4. Are always cheaper

9. Which observation most strongly signals genuine agentic adoption (vs AI theater) at a company?

  1. An AI mention buried in the annual report
  2. A chatbot on the website
  3. Workflows redesigned around agents
  4. An internal prompt-writing training course

10. The strategic risk of over-indexing on ONE flashy AI demo is…

  1. Demos are illegal to share
  2. Confusing a capability spike in a narrow case with reliable performance across your real distribution of work
  3. Demos expire
  4. Vendors dislike it

11. Revenue per employee is a useful inflection metric because…

  1. The HR team already carefully tracks it
  2. Output decouples from headcount
  3. It correlates closely with office size
  4. Regulators require it

12. Your industry analysis should end with…

  1. A neutral summary of all the various views
  2. A position on where value moves
  3. A long list of the possible vendors
  4. A disclaimer

Module 3 — The Leap to Agentic AI: From Generation to Action

Outcome: Distinguish generative AI, automation, and true agents — and ship your first working agentic workflow.

Leadership lens: The difference between "AI that answers" and "AI that acts" is the difference between a research analyst and a delegated employee. Leaders must know first-hand what delegation to software feels like — including the loss-of-control moment.

Apply-at-work mission — Automate one monitoring task: Pick one thing you or your team checks manually (inbox, dashboard, RSS, sheet) and build an n8n/Make agentic workflow that monitors it and takes an action (summarise + notify). Run it for real.

Reflection: Where did I feel a loss of control this week when the agent acted for me? Which controls restored my confidence — and which were just comfort theatre?

Resources

Project — First Agentic Workflow (Capstone Milestone 1)

Build your first real agentic workflow in n8n (or Make): an agent that monitors a source (email, RSS, form, sheet) and takes an action (summarise + notify, triage + route). Separately, run CrewAI's research-agent example and observe multi-agent collaboration. Then write a one-page comparison: what makes your workflow "agentic" vs a traditional Zapier-style automation — components (planner, memory, tools, executor, feedback) named explicitly. 🎯 This completes Capstone Milestone 1: a working agentic workflow you built yourself.

Deliverable: portfolio/w03-first-agent.md — workflow export/screenshots, what it does, the components map, and the automation-vs-agent comparison.

Assessment rubric

Criterion Weight What good looks like
It actually runs 30% The workflow executed on real inputs at least 5 times; you can show a run log and one output that reached you (or a colleague).
Component literacy 25% You can point at your workflow and name planner/memory/tools/executor/feedback loop — and honestly note which are missing.
Agentic vs automation analysis 25% The comparison identifies precisely where the LLM makes decisions vs where paths are fixed — no hand-waving about "smart automation".
Multi-agent observation 20% Notes from the CrewAI run: how agents divided the work, where they duplicated effort or drifted, and one thing that surprised you.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. The critical difference between generative AI and agentic AI is…

  1. Model size
  2. Generative AI produces content; agentic AI acts
  3. Generative AI is a much older technology overall
  4. Agentic AI is always multi-agent

2. The five core building blocks of an agent are…

  1. GPU, CPU, RAM, disk, and the network
  2. Planner, memory, tools, executor
  3. Prompt, response, rating, retry, cache
  4. Input, output, database, UI, and API

3. A Zapier automation that always runs the same fixed steps is…

  1. Actually a genuinely true agent
  2. A fixed workflow, not an agent
  3. A full multi-agent system of some kind
  4. Simply an LLM running on its own

4. Why do agents need a feedback/observation step in their loop?

  1. To bill more tokens
  2. To check results and re-plan when reality diverges
  3. Purely for logging and regulatory compliance purposes
  4. To slow them down safely

5. "Tools" in agent terminology means…

  1. Various developer IDEs and code editors
  2. Functions/APIs the agent can call
  3. Assorted physical hardware accessories
  4. Prompt templates

6. The safest first autonomy setting for a new workplace agent that can send messages is…

  1. Full autonomy, logs reviewed monthly
  2. Draft-and-approve first
  3. Read-only mode for ever more
  4. No logging at all, to keep it simple

7. In your n8n build, where does the "agentic" part actually live?

  1. In the trigger node
  2. In the LLM node(s) deciding how to categorise/summarise/route based on content
  3. In the notification step
  4. In the credentials store

8. Multi-agent collaboration (CrewAI-style) is best justified when…

  1. When a single prompt just gets far too long
  2. Distinct roles genuinely divide the work
  3. You simply want it to look more impressive
  4. Single agents are deprecated

9. An agent monitoring your inbox misfiles an important email. The leadership-grade response is…

  1. Shut down all agents
  2. Treat it as an incident: trace the decision, tighten the rule or gate, keep operating
  3. Ignore it as a one-off
  4. Blame the vendor

10. Real enterprise agent examples today (calendar, research, support triage) share what property?

  1. Complete full autonomy over the money
  2. Bounded domains with clear tools
  3. No memory at all
  4. They replace entire whole departments

11. The "loss of control" feeling when an agent acts for you is best handled by…

  1. Never delegating actions
  2. Designing observability and reversibility
  3. Only ever running the agents fully manually
  4. Trusting the vendor

12. Why does this programme make executives BUILD a workflow rather than just study agents?

  1. To convert them into engineers
  2. First-hand delegation-to-software experience changes how they lead, buy, and govern it
  3. Because reading is ineffective
  4. To fill the week

Phase 2 — Building: Vibe Coding & Agentic Workflows (weeks 4–8)

Module 4 — Vibe Coding Fundamentals

Outcome: Build your first functional app with natural language — and develop judgment for when vibe coding shines vs when it bites.

Leadership lens: You don't need to become an engineer — you need to stop being blocked by the absence of one. A leader who can prototype an idea before the next steering meeting changes the tempo of the whole organisation.

Apply-at-work mission — Ship a tool your team uses: Build one small internal tool with Claude Artifacts / Cursor (dashboard, notes processor, checklist app) and put it in front of at least two colleagues. Capture their reaction and one improvement request.

Reflection: What did it feel like to build something without "knowing how to code"? Where did vibe coding fail me, and what did that teach me about reviewing AI work I can't fully verify?

Resources

Project — Ship an Internal Tool by Describing It

Using only natural-language prompting (Claude Artifacts or Cursor), build a small but genuinely useful internal tool: a personal dashboard, meeting-notes processor, decision log, or content repurposer. Put it in front of at least two colleagues. Then attempt to significantly improve an existing small script or spreadsheet process by vibe coding, and write a short honest critique: "Where vibe coding failed me and what I had to fix manually."

Deliverable: portfolio/w04-vibe-tool.md — what you built, prompts that mattered, colleague feedback, and the failure critique.

Assessment rubric

Criterion Weight What good looks like
Usefulness 30% The tool addresses a real recurring annoyance; two colleagues used or reviewed it and one asked to keep it.
Prompt-to-product skill 25% You iterated: initial description → review → refinement cycles documented, showing you steered rather than accepted.
Critical judgment 30% The critique names specific failures (logic, edge cases, security assumptions) and what human judgment had to add — no cheerleading.
Risk awareness 15% You can state what this tool must NOT be used for (real data? external users?) and why.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. "Vibe coding" means…

  1. Coding with music on
  2. Describing intent in natural language and letting AI generate the implementation, which you review and refine
  3. Copying from Stack Overflow
  4. Programming without tests

2. The most important skill vibe coding does NOT remove is…

  1. Sheer raw typing speed at the keyboard
  2. Validating it does what you intended
  3. Knowing all of the keyboard shortcuts
  4. Memorising syntax

3. For a leader, the strategic value of being able to prototype is…

  1. Completely replacing the whole engineering team
  2. Changing tempo — ideas before the next meeting
  3. Saving licence fees
  4. Winning hackathons

4. Vibe-coded tools most dangerously accumulate…

  1. A large amount of unnecessary disk usage over time
  2. Hidden technical debt: unexamined edge cases
  3. Too many features
  4. Licensing costs

5. The right corporate policy for a useful vibe-coded prototype handling real customer data is…

  1. Ship it — it works
  2. Treat it as a validated idea: hand to engineering for a hardened rebuild before real data touches it
  3. Keep it secret from IT
  4. Add a password

6. When AI-generated code fails, it usually fails…

  1. Very loudly and really quite obviously
  2. Plausibly — fine on the happy path
  3. Only ever at the compile-time stage itself
  4. Entirely at random and unpredictably

7. Which request is vibe coding CURRENTLY best suited for?

  1. A core banking system ledger
  2. A personal dashboard tool
  3. Medical-device firmware code
  4. A high-frequency trading engine

8. The best prompting pattern for building beats one giant prompt because…

  1. Very long prompts always end up costing a lot more money
  2. Iterative describe → inspect → refine catches drift
  3. Models refuse long prompts
  4. Short prompts look professional

9. "I had to fix it manually" moments primarily teach leaders…

  1. That AI is basically completely useless
  2. Where AI competence ends today
  3. That they should probably code a lot more
  4. Nothing useful

10. A colleague proudly ships a vibe-coded customer-facing app with no engineering review. Your first question is…

  1. Which particular model did you use for it?
  2. Malicious input, and who maintains it?
  3. How many separate prompts did it take you?
  4. Can I have the prompt?

11. The "describing outcomes vs writing code" shift most resembles which leadership transition?

  1. Engineer → manager
  2. Intern → junior analyst
  3. Employee → startup founder
  4. Manager → the board member

12. Your vibe-coded tool works but you can't explain HOW. The leadership risk is…

  1. None — results matter
  2. You cannot reason about its failure modes, so you cannot responsibly decide where it may be used
  3. Colleagues will ask questions
  4. It might be plagiarised

Module 5 — Vibe Coding Advanced + API Integration

Outcome: Move beyond toys: build multi-step, API-connected applications that touch real business systems.

Leadership lens: The moment your prototype calls a real API, it stops being a demo and starts being a system — with error handling, credentials, and consequences. This is where executive prototypes earn (or lose) engineering's respect.

Apply-at-work mission — Connect a prototype to a real system: Extend a workflow or app so it calls at least one real external API (search, CRM, sheet, calendar) and processes the result with a second AI step. Document what you had to fix by hand.

Reflection: What architecture decisions did I make this week without realising they were architecture decisions? What would I now ask an engineering team that I couldn't have asked before?

Resources

Project — Multi-Step Research Agent with Real APIs

Build a multi-step agent that: (1) takes an input, (2) calls at least one real external API (web search, CRM, sheets, calendar), (3) processes results with a second AI step, and (4) outputs a structured action or report. The classic version: a personal research agent that searches the web and synthesises findings into a structured brief. Document your architecture decisions and everything you corrected by hand.

Deliverable: portfolio/w05-api-agent.md — architecture sketch, the working flow (screenshots/export), sample outputs, and the corrections log.

Assessment rubric

Criterion Weight What good looks like
End-to-end flow 30% Input → API call → AI processing → structured output all work on 5+ real runs, including one where the API returned something unexpected.
Architecture articulation 25% A simple diagram + prose naming each step, what can fail there, and what happens when it does.
Structured output 20% The final output follows a consistent, consumable format (fields, sections) — something a downstream system or colleague could rely on.
Corrections log 25% Honest record of what the AI got wrong and your fixes — evidence you reviewed rather than trusted.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. "Function calling" / tool use lets an LLM…

  1. Rewrite its own weights
  2. Emit a structured request that your system executes (API call, query), returning results to the model
  3. Call phone numbers
  4. Compile code faster

2. Why do multi-step agents need STRUCTURED outputs between steps?

  1. They simply look a good deal tidier overall
  2. Downstream steps parse fields, not prose
  3. The regulators specifically require the use of JSON
  4. It reduces tokens

3. An external API returns an error mid-flow. A well-designed agent…

  1. Just crashes very visibly indeed
  2. Retries, falls back, reports
  3. Invents plausible data to continue
  4. Ignores the step

4. The biggest NEW risk when your prototype connects to real business systems is…

  1. Slower demos
  2. Consequences: real writes and real data exposure
  3. Noticeably higher overall token bills every month
  4. More complex prompts

5. API credentials in an agent workflow should be…

  1. Simply pasted into the prompt
  2. In the credential vault
  3. Shared in team chat for convenience
  4. Hard-coded but lightly obfuscated

6. "Least privilege" for an agent's tools means…

  1. Simply using the very cheapest available API tier
  2. Granting only the permissions the task needs
  3. Running fewer steps
  4. One tool per agent

7. Chaining two AI steps (process API results with a second call) is useful because…

  1. Two separate models are pretty much always much smarter
  2. Each step gets a focused job, improving reliability
  3. It doubles the billing
  4. Single calls are deprecated

8. Your research agent cites a "fact" not present in any retrieved source. This is…

  1. Actually quite a useful bonus insight
  2. A grounding failure in synthesis
  3. Entirely normal and perfectly fine
  4. An API bug

9. Rate limits (429 responses) from an API should be handled by…

  1. Failing the entire whole run
  2. Backing off and retrying
  3. Switching to other vendors immediately
  4. Simply ignoring them entirely

10. Documenting architecture decisions matters for an executive builder because…

  1. It pads the portfolio
  2. It converts building experience into the vocabulary you'll use to govern engineering teams and vendors
  3. Engineers demand it
  4. It is required for the certificate

11. Which design makes an agent's work AUDITABLE?

  1. Deleting the logs to save some space
  2. Logging every step fully
  3. Using a much bigger, better model
  4. Password-protecting the whole workflow

12. The corrections log ("what I fixed by hand") is required because…

  1. Everyone makes mistakes
  2. It builds your calibration of where AI-generated systems need human verification — the core delegation skill
  3. It fills the report
  4. It shames the model

Module 6 — Designing Workflows for Agentic Systems

Outcome: Think computationally: map real business processes into agent-friendly workflows with clear decisions, exceptions, and human gates.

Leadership lens: Process mapping is old; mapping for agents is new. Every fuzzy handoff a human papers over becomes a failure mode when an agent runs it. The canvas you draw this week is the leadership artefact of the whole programme.

Apply-at-work mission — Canvas a real workflow with a colleague: Take one painful manual process from your organisation and complete the Agentic Workflow Canvas with the person who runs it: inputs, decisions, tools, outputs, exceptions, and the 5 places human judgment stays.

Reflection: Which steps of "my" processes could I actually not describe precisely when forced to? What does that say about how much of my organisation runs on tacit knowledge?

Resources

Project — Agentic Workflow Canvas (Capstone Milestone 2)

Take one painful manual process from your work or life and map it completely as an agentic workflow: inputs, decision points, tools, outputs, exception paths — and the 5 places human judgment must remain. Then design (on paper) a multi-agent system for a content or analysis pipeline (e.g. Researcher + Writer + Editor + Publisher) with a visual diagram and structured specification. 🎯 This completes Capstone Milestone 2: a real process, fully mapped for agents.

Deliverable: portfolio/w06-workflow-canvas.md + diagram — the completed canvas, the multi-agent spec, and the 5 human-judgment points.

Assessment rubric

Criterion Weight What good looks like
Process fidelity 25% Mapped with the person who actually runs the process (or your honest first-hand knowledge); includes the messy exceptions, not the idealised flow.
Decision & exception design 30% Every decision point has defined criteria; every exception has a path (retry, fallback, escalate) — nothing ends in "somehow".
Human-judgment placement 25% The 5 human gates are placed where stakes or ambiguity are genuinely high, with rationale — not sprinkled for comfort.
Multi-agent spec quality 20% Roles have distinct responsibilities and tools; handoffs specify what artifact passes between agents in what format.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. The three levels of workflow abstraction this week are…

  1. Code, config, and the docs
  2. Task → process → system
  3. Input, output, and storage
  4. Team, department, company

2. Computational thinking, for a non-engineer leader, chiefly means…

  1. Learning Python
  2. Decomposing fuzzy work into explicit steps, decisions, and data — precisely enough to delegate
  3. Doing mental arithmetic
  4. Thinking like a computer

3. Why do human workflows break when handed directly to agents?

  1. Agents are slow
  2. Human processes run on tacit knowledge
  3. The agents flatly refuse to do repetitive work
  4. APIs are unreliable

4. A decision point in an agent-ready workflow map must have…

  1. A single responsible senior executive
  2. Explicit criteria for each branch
  3. A hard timeout on the whole thing
  4. An AI model assigned

5. Exception paths deserve MORE design attention than happy paths because…

  1. They occur more often
  2. Failures cluster there, corrupting outcomes silently
  3. The auditors always read carefully through them first
  4. They are easier to design

6. Human-in-the-loop gates are best placed where…

  1. Wherever the work is at its most boring
  2. Stakes and ambiguity are high
  3. Wherever the agent happens to be fastest
  4. Wherever managers want more visibility

7. Designing for observability means…

  1. Making all the dashboards look pretty
  2. Each step can be inspected later
  3. Watching all the agents in real time
  4. Weekly status meetings

8. In multi-agent design, a "hierarchical" pattern means…

  1. The agents all vote on the decisions
  2. A supervisor delegates to workers
  3. All the agents just run in sequence
  4. All agents share one prompt

9. The most common failure of first-time workflow mappers is…

  1. Too many diagrams
  2. Mapping the idealised process, not the real one
  3. Simply using entirely the wrong tool for the job
  4. Too much detail

10. A good handoff spec between two agents defines…

  1. Their respective model versions
  2. The artifact and its format
  3. Their respective token budgets
  4. Which of them is more important

11. Why produce the canvas BEFORE building anything?

  1. Tools are expensive
  2. Design errors cost minutes on paper and weeks in production; the canvas is also your alignment artifact with stakeholders
  3. Building first is forbidden
  4. Diagrams impress leadership

12. You cannot precisely describe a step in "your" process when forced to. The leadership lesson is…

  1. The process is fine as folklore
  2. Your organisation runs on undocumented tacit knowledge — a risk AND the first thing agentification exposes
  3. Someone else should map it
  4. Skip that step

Module 7 — Automating & Orchestrating Workflows

Outcome: Orchestrate multiple agents and tools into a production-style system with logging, monitoring, and a runbook.

Leadership lens: One agent is a trick; an orchestrated system is an operating model. Sequential, parallel, hierarchical, swarm — these patterns are org charts for software teammates, and you're the one drawing them.

Apply-at-work mission — Stand up a multi-agent system: Build a multi-agent workflow that solves a real problem end-to-end (e.g. research → qualify → draft outreach), add basic logging, and write a one-page runbook: how it works, what can go wrong, who to call.

Reflection: If this system ran while I slept, what's the worst thing it could plausibly do? Did my design catch that — or did I only design for the happy path?

Resources

Project — Orchestrated Multi-Agent System with a Runbook

Build a complete multi-agent system solving a real problem end-to-end — the reference build: automated lead/topic research → qualification/analysis → prepared output (outreach draft, brief, or report). Implement basic logging/observability so you can reconstruct any run. Then write a real runbook: how it works, what can go wrong, how to tell, and what to do about it.

Deliverable: portfolio/w07-orchestrated-system.md — system description + diagram, run logs from 5+ real executions, and the runbook.

Assessment rubric

Criterion Weight What good looks like
End-to-end operation 30% The full chain runs on real inputs without manual patching between steps; at least 5 logged runs including one failure you can explain.
Orchestration pattern choice 20% You chose sequential/parallel/hierarchical deliberately and can defend why against one alternative.
Observability 25% Logs let you answer "what did it do and why" for any run without guessing; one failure diagnosed FROM the logs.
Runbook quality 25% A colleague could operate the system from the runbook alone: failure symptoms → diagnosis → response, including the kill switch.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. Piloting vs productionising an agentic workflow differ mainly in…

  1. The specific underlying model that gets used
  2. Error handling, monitoring, and ownership
  3. The prompt length
  4. The demo audience

2. Sequential orchestration is the right default when…

  1. You want the maximum possible speed
  2. Each step needs the last
  3. The agents disagree quite often
  4. You happen to have many GPUs

3. A "swarm" pattern (many agents, loose coordination) is…

  1. Always the most advanced choice
  2. Rarely justified — high coordination cost and failure surface for most business problems
  3. Required for production
  4. Cheaper than one agent

4. The purpose of logging every agent step is…

  1. Simply mere compliance box-ticking
  2. Reconstructing any run fully
  3. Slowing the whole system down safely
  4. Training data collection

5. A runbook exists so that…

  1. You remember your own system
  2. Someone who isn't the builder can operate it
  3. The auditors have some useful reading material
  4. The project looks professional

6. The single most important control in any autonomous workflow is…

  1. A much bigger model
  2. A fast kill switch
  3. A good many more agents
  4. Neatly colour-coded logs

7. Two of your agents produce contradictory outputs mid-pipeline. Good orchestration design…

  1. It simply just picks one of them at random
  2. Routes the conflict to a resolver
  3. It simply averages out all the answers
  4. Restarts everything

8. Cost runaway in multi-agent systems typically comes from…

  1. Expensive software licences
  2. Uncapped loops and retries
  3. Overly long variable naming
  4. Excessive logging overhead throughout

9. "It worked in my demo" fails in operation most often because…

  1. The demos simply use much better hardware
  2. Real inputs are messier and adversarial
  3. Users are less intelligent
  4. Networks are slower

10. Basic observability for a business-owned agent system minimally includes…

  1. The current GPU core temperature
  2. Status, traces, errors, cost
  3. Overall employee satisfaction levels
  4. Model weights

11. The leadership reason to build this system yourself (once) is…

  1. To replace your ops team
  2. To internalise the true effort, failure modes, and controls — calibrating how you buy, staff, and govern later
  3. To earn engineering credentials
  4. Because vendors are untrustworthy

12. "What's the worst thing this system could plausibly do while I sleep?" is a design question because…

  1. It motivates the team
  2. Autonomy means consequences happen without you present — so worst cases must be bounded by design, not attention
  3. It sounds impressive
  4. Regulations require insomnia

Module 8 — Scaling from Pilot to SOP

Outcome: Know why most AI pilots die — and design the path from experiment to pilot to operational to scaled.

Leadership lens: Every organisation has a pilot graveyard. The 4-level maturity model, redesigned SOPs, and human+agent KPIs are how you become the leader whose pilots ship instead of stall.

Apply-at-work mission — Maturity-assess a real pilot: Run the Maturity Assessor on one AI pilot in your organisation (yours or someone else's). Produce a one-page scaling plan: current level, blockers to the next level, redesigned SOP, and two human+agent KPIs.

Reflection: Which of my organisation's KPIs actively punish agentic ways of working? What would I measure instead if I trusted the system?

Resources

Project — Pilot-to-Scale Plan (Capstone Milestone 3)

Take one of your own workflows (week 3 or 7) or a real pilot from your organisation and build its scaling plan on the 4-level maturity model: Experiment → Pilot → Operational → Scaled. Redesign one existing SOP to include agentic components, and define KPIs that measure BOTH efficiency and quality/judgment in the human+agent process. 🎯 This completes Capstone Milestone 3: a credible path from pilot to operations.

Deliverable: portfolio/w08-scaling-plan.md — maturity assessment, blockers per level, the redesigned SOP, and the new KPI set.

Assessment rubric

Criterion Weight What good looks like
Honest maturity assessment 25% Current level justified with evidence; the temptation to grade your own pilot "operational" resisted.
Blocker analysis 25% Blockers to the next level are specific (ownership, error rate, integration, trust) with a named countermeasure each — not "needs more buy-in".
SOP redesign 25% The rewritten SOP specifies what the agent does, what humans do, and how exceptions and handoffs work — usable by the team tomorrow.
KPI design 25% At least two KPIs capture quality/judgment (not just speed/volume), each with a measurement method that would survive a sceptical CFO.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. The 4-level agentic maturity model runs…

  1. Idea → to demo → to launch → to exit
  2. Experiment → Pilot → Operational
  3. Crawl → then walk → then run → then fly
  4. Dev → to test → to stage → to prod

2. The most common reason AI pilots fail to scale is…

  1. Weak models
  2. No operating-model change: unclear ownership, unchanged processes, missing integration — the org, not the tech
  3. Insufficient GPUs
  4. Prompt quality

3. "Pilot purgatory" describes…

  1. A whole series of failed experiments
  2. Pilots that never reach operations
  3. A set of quietly cancelled projects
  4. Vendor lock-in

4. Moving from Pilot to Operational chiefly requires…

  1. A bigger model
  2. Named ownership, error-handling, and integration
  3. Signing up a great many more pilot users than before
  4. A press release

5. Traditional SOPs break in human+agent teams because…

  1. The agents simply cannot read the documents at all
  2. They assume a human fills the gaps at each step
  3. SOPs are obsolete generally
  4. Compliance forbids agents

6. A pure efficiency KPI ("tickets closed per hour") applied to an agentic process risks…

  1. Nothing — efficiency is the goal
  2. Rewarding fast wrong answers
  3. Somewhat slower overall adoption
  4. Union complaints

7. A good "quality/judgment" KPI for a human+agent workflow is…

  1. Total agent runs
  2. Escalation appropriateness or correction rate
  3. The total number of separate prompts written out
  4. Tokens consumed

8. Centralized vs federated workflow governance: the pragmatic answer for most orgs is…

  1. Fully centralized — one AI team owns everything
  2. Central standards, federated build
  3. Fully federated — every single team alone
  4. External vendors entirely own all of it

9. The "common failure pattern" of scaling a pilot everywhere immediately after one success is dangerous because…

  1. Success is repeatable by default
  2. One context's success hides context-specific conditions; scaled rollouts meet different data, users, and edge cases
  3. It is too cheap
  4. Legal forbids it

10. The right FUNDING model shift from pilot to scale is…

  1. A one-off project budget forever
  2. From project to product funding
  3. No funding at all — it's automated now
  4. Charge users per prompt

11. Your pilot's error rate is 8% with humans catching all errors. Before scaling, you must know…

  1. Nothing at all — humans catch them
  2. Whether review holds at scale
  3. The model's exact parameter count
  4. The competitors' own error rates

12. Redefining KPIs for agent-augmented teams is a LEADERSHIP task (not an analyst task) because…

  1. Analysts are busy
  2. KPIs encode what the organisation values; changing them changes behaviour, careers, and incentives — political territory
  3. Leaders like dashboards
  4. It requires no data

Phase 3 — Agentic AI Across the Business (weeks 9–11)

Module 9 — Agentic AI in Operations & Sales

Outcome: Design agent systems for pipeline, accounts, and operations — with compliance, audit trails, and human oversight built in.

Leadership lens: Ops and sales are where agentic ROI shows up first and where compliance failures show up loudest. Lead research agents, follow-up sequences, approval workflows — leverage with a paper trail.

Apply-at-work mission — Build a lead/ops research agent: Build a "research & qualification agent" for a real ops or sales motion: input a company/case, output a structured brief. Include one human approval gate and show it to whoever owns that pipeline.

Reflection: Where is the line in my business between "agent prepares, human decides" and "agent decides"? Who should own moving that line — and is it currently owned by anyone?

Resources

Project — Lead Research & Qualification Agent

Build a "Lead Research & Qualification Agent" for a real ops or sales motion: input a company name (or case), output a structured brief — ICP fit, key people, recent news, likely pain points. Add an automated follow-up-preparation step that respects compliance rules, and design a human-in-the-loop approval gate for any high-value action. Show it to whoever owns that pipeline.

Deliverable: portfolio/w09-sales-ops-agent.md — the working flow, 3 sample briefs, the compliance notes, and pipeline-owner feedback.

Assessment rubric

Criterion Weight What good looks like
Brief quality 30% Structured, sourced, decision-ready briefs a real seller/operator would use — verified against what they already know about one account.
Human approval gate 25% High-value actions demonstrably cannot proceed without human release; the gate shows WHAT the human approves, not just a yes button.
Compliance thinking 25% Data sources, consent constraints, and record-keeping named explicitly; nothing scraped or stored that the business couldn't defend.
Stakeholder validation 20% Feedback from the pipeline owner captured honestly, including what they would NOT trust the agent with.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. Ops and sales are typically the first agentic AI beachhead because…

  1. The salespeople all love technology
  2. High-volume, structured work
  3. They tend to have the biggest budgets
  4. Compliance is entirely absent there

2. A lead-research agent's output should be structured (ICP fit, people, news, pains) because…

  1. It looks professional
  2. Sellers act on decision-ready fields
  3. Free text is basically impossible to use
  4. CRMs demand JSON

3. Automated outreach WITHOUT human review risks…

  1. Only a few minor spelling errors
  2. Compliance and brand damage
  3. Somewhat slower overall pipelines
  4. Higher monthly API bills to pay

4. An audit trail in a sales/ops agent context means…

  1. A full backup copy of the entire CRM system
  2. A record of what was done and by whom
  3. All of the customer call recordings kept
  4. The prompt library

5. The human approval gate for high-value actions should show the approver…

  1. A progress bar
  2. The action, its evidence, and its consequences
  3. Just a simple single approve button for them to click
  4. The model temperature

6. "Pipeline intelligence" agents create most value by…

  1. Replacing account executives
  2. Compressing research and preparation time so humans spend their hours on judgment and relationships
  3. Sending more emails
  4. Automating discounts

7. Data used by a research agent must be checked for…

  1. Font compatibility
  2. Provenance and permission: is it public, licensed, or consented — and would you defend its use?
  3. File size
  4. Language

8. The agent confidently reports a "recent funding round" that never happened. The systemic fix is…

  1. Immediately fire the whole agent
  2. Ground claims in cited sources
  3. Simply lower the temperature setting
  4. Add a disclaimer

9. Which task should REMAIN human in an agent-augmented sales process?

  1. General company news gathering
  2. Final deal-strategy judgment
  3. Routine CRM field record updates
  4. General meeting scheduling work

10. Rolling this agent to the whole sales team after one rep's success requires first…

  1. Nothing at all — the success clearly proves it
  2. Checking the workflow fits other segments
  3. A bigger model
  4. A launch party

11. The pipeline-owner's "I wouldn't trust it with X" feedback is valuable because…

  1. It usefully identifies the sceptics
  2. It maps the trust boundary clearly
  3. It can be safely ignored after launch
  4. It fills the report

12. Measuring this agent's success should centre on…

  1. Number of briefs generated
  2. Downstream outcomes: qualified-meeting rate, time-to-first-touch, and correction rate of briefs
  3. Tokens used
  4. Rep satisfaction only

Module 10 — Agentic AI in Marketing & Customer Experience

Outcome: Build CX and content systems that scale personalisation without sacrificing brand, tone, empathy, or privacy.

Leadership lens: Marketing is the easiest place to deploy agents and the easiest place to damage a brand at machine speed. Brand-voice grounding, escalation paths, and consent-aware personalisation are leadership controls, not settings.

Apply-at-work mission — Repurpose content with brand guardrails: Build a content repurposing agent grounded in a real brand-voice document (yours or your employer's): one long-form piece → three platform-specific versions. Have the brand owner grade the outputs.

Reflection: What parts of my customers' experience should never be synthetic, even if no one could tell? Where do I draw that line and why?

Resources

Project — Brand-Safe Content & CX Agents

Two builds. (1) A content repurposing agent grounded in a real brand-voice document: one long-form piece in, three platform-specific versions out — graded by the brand owner. (2) A customer-support triage agent design (build if time allows): categorise incoming requests, draft responses for routine cases, and define clear escalation paths to humans. Plus: a one-page personalisation policy covering privacy, consent, and tone.

Deliverable: portfolio/w10-brand-cx-agents.md — repurposer outputs + brand-owner grades, the triage design, and the personalisation policy.

Assessment rubric

Criterion Weight What good looks like
Brand-voice fidelity 30% The brand owner grades outputs ≥7/10 on voice; deviations analysed — you know WHY it drifted where it did.
Platform adaptation 20% The three versions genuinely differ in form and register for their platforms, not just in length.
Escalation design 25% Triage rules name the categories agents may answer, must draft-only, and must hand straight to humans — with the signals that trigger each.
Personalisation policy 25% Privacy/consent/tone rules concrete enough to enforce: what data may personalise what, and what must never be synthetic.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. The core tension in agentic marketing is…

  1. Overall cost versus raw speed
  2. Scale vs brand consistency
  3. SEO versus paid online advertising
  4. Plain text versus rich video

2. Grounding a content agent in a brand-voice document works because…

  1. It shortens the prompts you need to write out
  2. It generates against explicit constraints
  3. It reduces cost
  4. Legal requires it

3. "Human oversight" in agentic content creation is best implemented as…

  1. Reading everything after publication
  2. Review gates by risk tier
  3. Trusting the prompt
  4. Weekly audits only

4. A support triage agent should hand to a human IMMEDIATELY when…

  1. The queue happens to be long
  2. It detects distress or risk
  3. The customer writes in all caps
  4. After about three or so exchanges

5. Customer journey orchestration with agents means…

  1. Simply sending more emails per customer
  2. Coordinating touchpoints coherently
  3. Automating all of the churn surveys sent
  4. A/B testing everything

6. Consent-aware personalisation requires…

  1. Personalising everything that is possible
  2. Only data provided for that purpose
  3. Anonymous background tracking of them
  4. Longer privacy policies

7. The biggest CX risk of routine-response automation is…

  1. Occasional minor grammar mistakes here
  2. Confidently wrong replies at scale
  3. Somewhat slower responses overall now
  4. Higher costs

8. "What should never be synthetic" is a leadership question because…

  1. Engineers can't decide fonts
  2. It draws the brand's authenticity line — a values decision with commercial consequences, not a technical setting
  3. Regulators publish the list
  4. It changes weekly

9. Measuring a repurposing agent purely on output volume incentivises…

  1. Genuinely better quality
  2. Off-brand content
  3. Much better research
  4. Nothing at all — volume is neutral

10. Brand-owner grading of agent outputs (your project) establishes…

  1. Who is boss
  2. A quality baseline and a feedback signal
  3. A useful published marketing case study for it
  4. Legal cover

11. An agent personalises a message using data the customer never knowingly shared. Even if legal, this is…

  1. A win for relevance
  2. A trust liability: perceived surveillance damages the relationship personalisation was meant to deepen
  3. Standard practice — fine
  4. A/B testable

12. Escalation paths must be designed BEFORE deploying CX agents because…

  1. Documentation looks good
  2. Under load, undefined escalation collapses into either everything-to-humans or nothing-to-humans
  3. Agents demand them
  4. They cannot be changed later

Module 11 — Agentic AI in HR & Finance

Outcome: Apply agents to people and money decisions with bias mitigation, auditability, and accountability that would survive a regulator.

Leadership lens: HR and Finance are high-leverage AND high-risk: résumé screening, anomaly detection, performance prep. The design question is never "can the agent do it?" — it's "who is accountable when it does?"

Apply-at-work mission — Design a review-gated HR/Finance workflow: Design (build if you can) one HR or Finance agent workflow with explicit bias-mitigation steps and human review gates — e.g. screening assistant or spend-anomaly reporter. Write down its audit trail.

Reflection: If an agent-assisted decision about a person turned out wrong, could I explain the decision chain to that person's face? What would need to change so I could?

Resources

Project — Accountable HR/Finance Agent Workflows

Design — and build what you can — two sensitive-domain workflows: (1) a résumé screening + initial outreach assistant with explicit bias-mitigation steps and human review gates; (2) a financial anomaly detection + reporting agent (n8n + sheets + LLM analysis works). For each, document the audit trail: what is recorded, who is accountable for each decision, and how a challenged decision would be reconstructed and explained.

Deliverable: portfolio/w11-hr-finance-agents.md — both designs, the bias-mitigation measures, audit-trail specs, and an accountability map.

Assessment rubric

Criterion Weight What good looks like
Bias mitigation 30% Concrete measures (structured criteria before screening, blind fields, sampled human audits of rejections) — not a statement that bias is bad.
Human accountability 25% Every consequential decision has a named human owner; the agent recommends, a person decides — and the map shows it.
Audit trail design 25% A challenged decision (rejected candidate, flagged transaction) can be fully reconstructed: inputs, criteria, model output, human action.
Risk/leverage judgment 20% The write-up distinguishes where automation is high-leverage (drudgery) vs high-risk (judgment about people/money) with defensible placements.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. HR and Finance are "high-risk, high-leverage" domains for agents because…

  1. They happen to have by far the very most data
  2. Huge drudgery, but errors touch livelihoods
  3. Budgets are centralised
  4. They adopt slowly

2. The correct division of labour for résumé screening is…

  1. The agent decides, human is informed
  2. Agent summarises; human decides
  3. Agent decides, with an appeal option
  4. Human does everything

3. AI screening can AMPLIFY hiring bias because…

  1. Models dislike certain fonts
  2. It learns patterns from historical decisions — including the biased ones — and applies them consistently at scale
  3. It reads too fast
  4. Candidates game it

4. A meaningful bias-mitigation step is…

  1. Telling the model to be fair
  2. Defining structured job-relevant criteria BEFORE screening and auditing a sample of rejections against them
  3. Using a larger model
  4. Removing all human review

5. For a financial anomaly-detection agent, a false NEGATIVE (missed anomaly) vs false POSITIVE (false alarm) trade-off should be set by…

  1. Simply the model's own built-in defaults
  2. Finance leadership's risk appetite
  3. Whatever minimises the alerts sent
  4. The vendor

6. "Auditability" in these domains concretely means…

  1. Annual audits happen
  2. Any decision can be reconstructed later
  3. The CFO has a nice live dashboard for it
  4. Logs exist somewhere

7. An agent flags an employee expense as anomalous. The next step should be…

  1. An automatic payroll deduction
  2. A human reviews it first
  3. An automated warning email sent
  4. Just silent logging of it

8. Explaining an agent-assisted decision "to the person's face" is the right design test because…

  1. It is emotionally satisfying
  2. It forces reconstructable, criteria-based decisions — if you can't explain it, you shouldn't operate it
  3. Lawyers recommend it
  4. It shortens meetings

9. Which HR task is the SAFEST early agent deployment?

  1. Making termination decisions
  2. Drafting role descriptions
  3. Setting performance ratings
  4. Setting people's actual salaries

10. "The agent recommends, the human decides" fails in practice when…

  1. The agents are simply too slow
  2. Humans rubber-stamp it
  3. The recommendations are too good
  4. The data is all well-formatted

11. Tracking how often human reviewers DISAGREE with agent recommendations is valuable because…

  1. It usefully ranks all of the employees
  2. Near-zero disagreement is a warning
  3. It quietly trains up the model further
  4. It fills dashboards

12. Regulatory exposure for AI in HR/Finance (GDPR/AI-Act-style rules) most concerns…

  1. Model size limits
  2. Automated decisions about individuals
  3. Open-source software licensing terms only
  4. Cloud regions only

Phase 4 — Leading in the Agentic Age (weeks 12–15)

Module 12 — Judgment in an Age of Abundance

Outcome: Intelligence is abundant; judgment is scarce. Learn to lead with WHY–WHAT–HOW when execution is nearly free.

Leadership lens: When AI can execute anything in minutes, the bottleneck moves to whoever decides what's worth executing. Decision velocity and learning velocity become the new leadership metrics — framing beats doing.

Apply-at-work mission — Run a WHY–WHAT–HOW conversation: Apply the WHY–WHAT–HOW framework to one live strategic question and use it to structure a real 15-minute conversation with your manager or team. Note where the framework changed the outcome.

Reflection: Write the first draft of your leadership manifesto: how will I lead differently in an age of abundant intelligence? Which of my current strengths become commodities?

Resources

Project — Leadership Manifesto + WHY–WHAT–HOW in Practice

Deep-thinking week. (1) Write your personal leadership manifesto: "How I will lead differently in an age of abundant intelligence" — using an AI as a Socratic sparring partner that challenges rather than agrees. (2) Apply the WHY–WHAT–HOW framework to one live strategic challenge in your organisation and document the result. (3) Identify 3 areas where FASTER AI execution would create problems without better human framing.

Deliverable: portfolio/w12-manifesto.md — the manifesto, the WHY–WHAT–HOW application, and the 3 speed-traps.

Assessment rubric

Criterion Weight What good looks like
Manifesto depth 30% Names specific behaviours you will start/stop/change — commitments a colleague could hold you to, not aspirational fog.
Socratic process 20% Evidence the AI challenged you: at least two positions you revised under questioning, documented.
Framework application 30% The WHY–WHAT–HOW analysis of a real challenge shows the layers separated cleanly — and reveals something the default framing missed.
Speed-trap insight 20% The 3 areas are real and defensible: places where execution velocity would amplify a framing error.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. The "AI paradox" of execution states…

  1. The AI simply cannot execute anything at all
  2. Direction is the binding constraint
  3. Execution no longer really matters much
  4. Paradoxes sell books

2. In the WHY–WHAT–HOW framework, leaders should increasingly concentrate on…

  1. HOW — implementation detail
  2. WHY and WHAT — purpose and framing
  3. None of them; all three equally forever
  4. Whichever is urgent

3. "Decision velocity" as a leadership metric means…

  1. Deciding everything instantly
  2. The speed from information to committed, high-quality decision — matched to the new pace of execution
  3. Number of decisions per week
  4. Delegating all decisions

4. "Learning velocity" matters more than stock of knowledge now because…

  1. Knowledge is expensive
  2. Capability and context shift so fast that the rate of updating beats the depth of any static expertise
  3. Universities are slow
  4. Memory fades

5. Moving from "executor" to "framer and orchestrator" means…

  1. Doing less work
  2. Defining problems, constraints, and quality bars
  3. Simply attending a good many more meetings than before
  4. Approving others' work

6. A team ships a polished, wrong deliverable in record time. The root failure is most likely…

  1. A team of overly lazy agents
  2. Bad framing of the WHAT
  3. Simply insufficient compute
  4. Just plain bad luck really

7. Using AI as a SOCRATIC partner (vs an answer engine) means…

  1. Simply asking rather harder questions of it
  2. Having it interrogate your assumptions
  3. Using two entirely separate models at once
  4. Asking twice

8. Which decision still deserves SLOW deliberation despite fast execution being available?

  1. Reformatting a whole report
  2. An irreversible commitment
  3. Choosing everyone's meeting times
  4. Renaming a single small file

9. "Intelligence is abundant, judgment is scarce" implies organisations should now compete on…

  1. Raw underlying model access rights
  2. Problem selection and framing
  3. Big shared prompt libraries kept
  4. Compute supply contracts signed

10. The most dangerous leadership response to abundant execution capacity is…

  1. Careful, deliberate prioritisation
  2. Filling capacity with more
  3. Training up the whole team properly
  4. Measuring outcomes

11. A leadership manifesto beats vague intentions because…

  1. It reads well
  2. Written, specific commitments enable self-accountability and let others call the gap between stated and lived
  3. HR requires it
  4. It is shareable

12. Framing work ("what problem are we actually solving?") resists automation because…

  1. The models are generally quite bad at using words
  2. It needs context, stakes, and values judgment
  3. It is low-status work
  4. It changes rarely

Module 13 — The Three Roles That Scale: Orchestrator, Architect, Steward

Outcome: Master the three roles that gain value in agentic organisations — and diagnose which one you and your team are missing.

Leadership lens: Traditional management scales by headcount; agentic organisations scale by orchestration, architecture, and stewardship. Knowing which role you naturally play — and which your team lacks — is a career-defining insight.

Apply-at-work mission — Role-map your team: Map yourself and your team against Orchestrator / Architect / Steward. Identify the missing role, then draft a job description for an "Agentic AI Orchestrator" (or the role you lack) in your organisation's language.

Reflection: Which of the three roles do I gravitate to under pressure — and is that the role my organisation actually needs most from me right now?

Resources

Project — Role Map & Agentic Team Blueprint (Capstone Milestone 4)

Map your current role and team against the Orchestrator–Architect–Steward model: who covers what, where the gaps are, which role you gravitate to. Design a small team structure (3–5 people + agents) optimised around one primary role for a real objective. Write a serious job description for an "Agentic AI Orchestrator" (or your organisation's missing role) in your company's language. 🎯 This completes Capstone Milestone 4: an organisational design you could actually propose.

Deliverable: portfolio/w13-role-blueprint.md — the role map with gaps, team design, and the job description.

Assessment rubric

Criterion Weight What good looks like
Honest role mapping 25% Real names/functions mapped with evidence; the missing role identified from observed failures, not theory.
Team design coherence 30% The 3–5 person + agents structure names each human's role, each agent's scope, and the decision rights between them.
Job description quality 25% The JD would survive HR review: responsibilities, competencies, success measures at 6/12 months — in your org's idiom.
Self-awareness 20% Your own gravitational role identified with supporting evidence, plus a development plan for your weakest of the three.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. The three roles that scale in agentic organisations are…

  1. The coder, the tester, the deployer
  2. Orchestrator, Architect, Steward
  3. The manager, director, and the VP
  4. The analyst, designer, and marketer

2. The Orchestrator's core skill is…

  1. Writing prompts
  2. Aligning humans and agents across silos — sequencing, handoffs, and shared context toward an outcome
  3. Approving budgets
  4. Debugging code

3. The Architect's core skill is…

  1. Ongoing people and team-management work
  2. Designing workflows and decision logic
  3. Model training
  4. Vendor negotiation

4. The Steward's core skill is…

  1. Aggressive cost cutting throughout
  2. Ensuring trust and accountability
  3. Loudly marketing all of the AI wins
  4. Writing policies nobody reads

5. These roles scale differently from traditional management because…

  1. They generally tend to pay a lot more
  2. Leverage grows with the workflows
  3. They genuinely need no real skills
  4. They avoid meetings

6. A team ships agent workflows fast but keeps having governance incidents. The missing role is…

  1. Orchestrator
  2. Steward
  3. Architect
  4. More engineers

7. Workflows exist and are governed, but nothing connects across departments and agents duplicate work. Missing role…

  1. Steward
  2. Orchestrator
  3. Architect
  4. Project manager

8. Everyone uses AI enthusiastically but workflows are ad-hoc, fragile, and undocumented. Missing role…

  1. Orchestrator
  2. Architect
  3. Steward
  4. Trainer

9. Middle management's traditional information-relay function is threatened because…

  1. The executives all got much friendlier
  2. Agents relay information instantly
  3. Most offices simply went fully remote
  4. Budgets shrank

10. A serious "Agentic AI Orchestrator" JD should measure success by…

  1. Number of prompts written
  2. Cross-functional workflow outcomes: cycle time, quality, and adoption of orchestrated human+agent processes
  3. Meetings hosted
  4. Certifications earned

11. Knowing which role YOU gravitate toward matters because…

  1. It largely determines your salary
  2. You over-supply your own role
  3. The roles are all quite permanent
  4. The HR team always tends to ask

12. A 4-person team + agents "optimised for one primary role" means…

  1. Everyone does everything
  2. The team's structure, rituals, and metrics centre on that role's leverage — e.g. an architecture pod that designs workflows for many teams
  3. Three people are idle
  4. One person rules

Module 14 — Cognitive Capital vs Cognitive Debt

Outcome: Design personal and team practices where AI amplifies thinking instead of replacing it — before automation bias compounds.

Leadership lens: Every convenience is a loan against a skill. Teams that offload thinking accrue cognitive debt that comes due at the worst moment — in a crisis, when the AI is wrong and nobody can tell. Leaders set the repayment schedule.

Apply-at-work mission — Audit your AI usage: Run the Cognitive Capital Audit on your own last two weeks of AI usage: where did AI deepen your thinking, where did it replace it? Adopt one "Thinking Amplification Protocol" and practice it for 5 days.

Reflection: Which thinking skill have I quietly stopped practicing since I started using AI daily? Do I want it back — and what's my plan if the answer is yes?

Resources

Project — Cognitive Capital Audit + Thinking Amplification Protocol

Audit your own AI usage over the past two weeks: for each significant use, classify — did AI deepen your thinking (capital) or replace it (debt)? Identify 3 areas of accumulating cognitive debt. Design your personal "Thinking Amplification Protocol" — a workflow where AI deepens understanding instead of shortcutting it — and practice AI Socratic dialogue on one complex topic: the AI challenges your assumptions rather than answering.

Deliverable: portfolio/w14-cognitive-audit.md — the usage audit, 3 debt areas, your protocol, and the Socratic dialogue transcript with commentary.

Assessment rubric

Criterion Weight What good looks like
Audit honesty 30% Real usage examined without flattering yourself; at least one uncomfortable finding about your own offloading.
Debt diagnosis 20% The 3 debt areas name specific skills degrading (structuring arguments, estimation, first-draft thinking) with evidence.
Protocol design 30% The protocol is concrete and practiced for 5 days: think first → AI challenges → revise — not an aspiration but a routine with a trigger.
Socratic transcript 20% The dialogue shows genuine challenge: assumptions surfaced, at least one position revised, commentary on what the process felt like.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. Automation bias is…

  1. Simply preferring newer tools
  2. Over-trusting the machine
  3. An outright hatred of manual work
  4. A kind of model training error

2. Cognitive debt accumulates when…

  1. You think too hard
  2. Skills atrophy because AI does them for you
  3. Your working memory quietly fills right up now
  4. You take notes

3. Cognitive capital, by contrast, is built when…

  1. You avoid using AI entirely and completely
  2. AI use deepens your understanding
  3. You save prompts
  4. You read more

4. The clearest warning sign of personal cognitive debt is…

  1. Simply using AI every single day
  2. Being unable to judge the work
  3. Generally preferring dictation now
  4. Typing slower

5. "Think first, then consult AI" beats "ask AI first" for important problems because…

  1. It is more polite
  2. Forming your own position first preserves the thinking skill and makes you a critic rather than a consumer of the AI's answer
  3. AI answers improve later in the day
  4. It saves tokens

6. For a TEAM, cognitive debt shows up as…

  1. Consistently much longer meetings
  2. Nobody can explain the decisions
  3. A great deal more documentation produced
  4. Slower hiring

7. Metacognition in this context means…

  1. Thinking quickly
  2. Monitoring your own thinking closely
  3. Various memory and recall techniques used
  4. Meditation

8. A Thinking Amplification Protocol should REQUIRE…

  1. Avoiding AI for the hard problems
  2. Your own draft before AI
  3. Using two separate AI models
  4. Longer prompts

9. Leaders bear special responsibility for cognitive debt because…

  1. They personally use AI the very most
  2. They set the norms teams copy
  3. Their own skills matter far less
  4. They approve all of the tools used

10. Which AI use pattern builds capital rather than debt for a strategy question?

  1. "Write our 3-year strategy"
  2. "Here is my draft strategy and reasoning — find the three weakest assumptions and argue against them"
  3. "Summarise our industry"
  4. "What would McKinsey say?"

11. The crisis-scenario argument for maintaining thinking skills is…

  1. Crises are rare
  2. In novel, high-stakes moments AI is least reliable and human judgment most needed — exactly when debt comes due
  3. Regulations require it
  4. Boards prefer it

12. "Every convenience is a loan against a skill" implies the leadership discipline of…

  1. Simply refusing all convenience
  2. Choosing which skills to keep
  3. Banning AI use for all juniors
  4. Regular weekly memory tests for all

Module 15 — The AI Paradox: Adoption Without Transformation

Outcome: Diagnose why high AI adoption so often yields low transformation — and design change strategies that beat organisational antibodies.

Leadership lens: "Everyone uses Copilot" and "nothing has changed" are both true in most enterprises. AI theater, structural mismatch, and middle-management antibodies are diagnosable and treatable — if a leader is willing to name them.

Apply-at-work mission — Run an adoption audit: Conduct an honest AI Adoption Audit of your team or organisation: where is adoption high but transformation low? Name the top 3 organisational antibodies and one countermeasure for each.

Reflection: Where am I personally performing AI theater — visible adoption without changed outcomes? What would real transformation of my own role look like?

Resources

Project — AI Adoption Audit + Change Strategy

Conduct an honest AI Adoption Audit of your team or organisation: where is adoption high (tools bought, accounts active) but transformation low (workflows, structures, and outcomes unchanged)? Identify the top 3 "organisational antibodies" resisting deeper agentic adoption in your context — and design a change strategy addressing both technical capability and organisational dynamics, with one countermeasure per antibody.

Deliverable: portfolio/w15-adoption-audit.md — the audit findings, 3 antibodies with evidence, and the change strategy.

Assessment rubric

Criterion Weight What good looks like
Audit rigour 30% Adoption vs transformation distinguished with observable evidence (usage stats vs changed processes/KPIs), not impressions.
Antibody diagnosis 25% The 3 antibodies are specific to your organisation (e.g. utilisation-based billing, review bottlenecks, role-protection in layer X) — with the incentive behind each named.
Change strategy realism 30% Countermeasures address incentives and structure, not just communication; each has an owner, a first step, and a way to tell it's working.
Self-inclusion 15% Your own AI theater identified — where your visible adoption exceeds your changed outcomes.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. The "AI paradox" of adoption without transformation describes…

  1. Nobody in the whole company using AI
  2. High usage, unchanged outcomes
  3. AI quietly replacing every single person
  4. A totally failed procurement process

2. "Organisational antibodies" are…

  1. A kind of security software
  2. Things that block change
  3. Certain difficult employees
  4. The various compliance rules

3. A classic structural mismatch blocking transformation is…

  1. The offices are simply far too small for all of it
  2. Human-speed structures around machine-speed work
  3. Too few licences
  4. Wrong cloud region

4. Middle-management resistance to agentic transformation is usually driven by…

  1. Widespread technophobia among the managers
  2. Rational role protection by managers
  3. Certain age demographics being at play
  4. Union rules

5. "AI theater" is…

  1. Corporate training videos
  2. Visible AI activity only
  3. AI used in the entertainment industry
  4. Long vendor presentations

6. The most reliable test distinguishing transformation from adoption is…

  1. Survey sentiment
  2. Whether workflows, decision rights, or KPIs have structurally changed — could the old process still run unchanged?
  3. Licence utilisation
  4. Executive quotes

7. Why do communication-only change programmes fail against antibodies?

  1. Emails go unread
  2. Antibodies live in incentives and structures; persuasion doesn't change what people are paid and promoted for
  3. Wrong channels
  4. Too few townhalls

8. A countermeasure for "utilisation-billed teams resist efficiency gains" is…

  1. More training
  2. Change the incentive model: price outcomes
  3. Simply mandate all AI usage across the board
  4. Hire consultants

9. Real capability building (vs theater) looks like…

  1. A regular internal company AI newsletter
  2. Redesigned workflows in production
  3. A big internal prompt-writing competition
  4. An innovation lab tour

10. Including YOURSELF in the adoption audit matters because…

  1. Humility is fashionable
  2. Leaders' own theater licenses everyone else's; credibility to demand change requires having made it personally
  3. Auditors require it
  4. It softens findings

11. The correct sequencing for beating antibodies is usually…

  1. A big company-wide mandate right on the first day
  2. Prove value in one real workflow, then expand
  3. Wait for competitors
  4. Reorg first

12. High adoption with low transformation is DANGEROUS (not merely disappointing) because…

  1. The software licences all expire
  2. It inoculates the organisation
  3. The vendors all get rich from it
  4. IT gets blamed

Phase 5 — Responsible Scale & Capstone (weeks 16–18)

Module 16 — Responsible AI and Agentic Scale

Outcome: Build governance for systems that act: guardrails + monitoring over approval gates, with trust, recourse, and rollback at scale.

Leadership lens: A hallucination in a chatbot is an embarrassment; a hallucination in an agent with system access is an incident. Governance design — autonomy tiers, monitoring, escalation, rollback — is the steward's core craft.

Apply-at-work mission — Draft a real risk register: Take one live or proposed agentic workflow in your organisation and produce a Risk Register + governance one-pager: failure modes, autonomy tiers, monitoring signals, escalation path, rollback plan.

Reflection: What's the riskiest thing my organisation currently lets AI do with no monitoring? Why has nobody asked — and what does it cost me to be the one who asks?

Resources

Project — Governance Framework + Risk Register (Capstone Milestone 5)

Design a governance framework for agentic workflows in your organisation: autonomy tiers (what agents may do freely / with approval / never), monitoring signals, escalation paths, and rollback procedures. Build a Risk Register for one real or proposed agentic system covering hallucination, action, and compliance risks — each with likelihood, impact, and a named control. If you can, add a simple monitoring/logging view for one of your own workflows. 🎯 This completes Capstone Milestone 5: governance you could table at a risk committee.

Deliverable: portfolio/w16-governance.md — the framework one-pager, the risk register, and (optional) monitoring screenshots.

Assessment rubric

Criterion Weight What good looks like
Autonomy tier design 25% Three-plus tiers with concrete examples per tier from your context; tier boundaries justified by reversibility and stakes.
Risk register quality 30% 8+ real risks across hallucination/action/compliance, each with likelihood, impact, owner, and a control that would actually work.
Monitoring & escalation 25% Named signals (error rate, cost, drift, complaint rate) with thresholds, plus who gets paged and what they do.
Rollback realism 20% A stop-and-recover procedure that works at 2am without the builder: kill switch, fallback process, communication plan.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. Agentic systems carry categorically higher risk than chatbots because…

  1. They cost more
  2. Hallucination plus action: errors become side effects
  3. They tend to run more or less continuously all the time
  4. They use more tokens

2. Autonomy tiers structure governance by…

  1. Ranking agents by IQ
  2. Defining what actions run freely, which need approval, and which are forbidden — by stakes and reversibility
  3. Pricing model access
  4. Separating dev and prod

3. The shift "from approval gates to guardrails + monitoring" happens because…

  1. Approvals are expensive
  2. At scale, per-action human approval collapses; continuous constraints + observability keep control without the bottleneck
  3. Regulators changed rules
  4. Guardrails sound better

4. A good monitoring SIGNAL for an agentic workflow is…

  1. The overall server uptime figure alone
  2. Behavioural rates, not just uptime
  3. The total number of active users only
  4. Model version

5. A risk register entry is complete when it has…

  1. A suitably scary-sounding name
  2. An owner and a control
  3. A neat colour-coded label
  4. A full insurance quote for it

6. "Recourse" as a trust property means…

  1. A particular form of legal insurance
  2. Affected parties can reach a human
  3. Standard company refund policies in place
  4. Apology templates

7. The rollback plan must work "at 2am without the builder" because…

  1. Builders sleep late
  2. Autonomous systems fail on their own schedule
  3. The auditors always run all their tests at night
  4. It sounds rigorous

8. Portfolio governance for agentic initiatives means…

  1. One mega-project
  2. Managing the whole set with shared standards
  3. A single steering committee with its own logo
  4. Annual reviews

9. Transparency about agent involvement (disclosure) matters because…

  1. The marketing team generally prefers it
  2. Discovered concealment destroys trust
  3. It is always strictly legally required
  4. Users read policies

10. Which failure is an ACTION risk (vs hallucination risk)?

  1. Inventing a statistic in a summary
  2. A wrongly sent refund
  3. Citing the wrong page number
  4. Overly verbose answers

11. An agent's behaviour slowly drifts as usage patterns change. The governance control is…

  1. Just simply hoping for the best
  2. Baseline metrics and alerts
  3. Restarting the whole thing weekly
  4. Using much bigger prompts throughout

12. Presenting governance as an ENABLER (not a brake) is credible when…

  1. It has a nice logo
  2. Clear tiers and pre-approved guardrails let teams ship low-risk automation FASTER, with escalation only where stakes demand
  3. It is optional
  4. Fines are mentioned

Module 17 — Continuous Transformation: Leading When AI Never Stops Evolving

Outcome: Build the operating rhythm — funding, teams, iteration cadence — for an organisation where the capability curve never flattens.

Leadership lens: This transformation has no "done". AI factories, agent-to-agent ecosystems, persistent iteration teams — leaders must normalise permanent beta without exhausting their people. That's an operating-model design problem.

Apply-at-work mission — Pitch an iteration team: Draft and pitch (to a real stakeholder, even informally) a lightweight funding + governance model for a persistent AI iteration capability in your organisation: who, budget envelope, cadence, kill criteria.

Reflection: What's my personal operating rhythm for staying current without drowning? What did I stop doing to make room for it?

Resources

Project — Continuous Transformation Playbook

Write the "Continuous Transformation Playbook" for your team or organisation: operating rhythm for ongoing AI iteration, a lightweight funding + governance model for a persistent iteration capability (who, budget envelope, cadence, kill criteria), and 3 scenarios for how agentic AI reshapes your industry/function in 2028–2030 — each with a "we would start doing X now" implication. Pitch the iteration-team model to a real stakeholder, even informally, and record the reaction.

Deliverable: portfolio/w17-transformation-playbook.md — the playbook, funding model, 3 scenarios, and the pitch outcome.

Assessment rubric

Criterion Weight What good looks like
Operating rhythm 25% A concrete cadence (weekly/monthly/quarterly loops) naming what gets evaluated, adopted, retired — sustainable, not heroic.
Funding & governance model 30% A real proposal: team shape, budget envelope, decision rights, and explicit kill criteria for experiments.
Scenario quality 25% Three genuinely different 2028–2030 futures with present-tense implications — each names something to START now.
Real pitch 20% You pitched a real person; their objections and your revisions documented.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. AI transformation differs from ERP/cloud transformations because…

  1. It is generally cheaper
  2. It has no end state
  3. It needs no real training
  4. The vendors are all nicer

2. A persistent "AI iteration team" exists to…

  1. Write out all of the prompts for everyone else here too
  2. Continuously evaluate, upgrade, and retire workflows
  3. Run the helpdesk
  4. Negotiate licences

3. Kill criteria for AI experiments matter because…

  1. The legal team strictly demands them
  2. Zombie pilots drain the budget
  3. They really do motivate the teams
  4. The boards all rather like the word

4. "Normalising iteration" for live agentic systems means…

  1. Holding weekly all-hands meetings
  2. Treating updates as routine
  3. Freezing all the systems yearly
  4. Blaming users less

5. Funding persistent iteration as OPEX/product (vs one-off project CAPEX) matters because…

  1. The accounting department simply always prefers it
  2. Project money ends and the capability decays
  3. It is cheaper
  4. CFOs never ask about OPEX

6. "AI factories" as a future trend refers to…

  1. Chip fabrication plants
  2. Organisations industrialising the production of AI-powered workflows — pipelines that turn processes into agentic systems repeatably
  3. Robot warehouses
  4. Prompt marketplaces

7. Agent-to-agent (A2A) protocols would strategically matter because…

  1. They cut API costs
  2. Your agents will transact with suppliers' and customers' agents — interfaces, trust, and standards become competitive terrain
  3. They are faster
  4. They simplify logging

8. Scenario planning beats point predictions for 2028–2030 because…

  1. Predictions are boring
  2. Under deep uncertainty, preparing for several divergent futures builds capabilities that pay off across most of them
  3. Scenarios need no data
  4. Consultants sell them

9. The biggest HUMAN risk of permanent transformation is…

  1. General wage and salary inflation
  2. Change fatigue and burnout
  3. Simply far too many meetings
  4. An overall surplus of skills

10. A "learning organisation" in the agentic era is distinguished by…

  1. A very big annual training budget
  2. Short feedback loops, in weeks
  3. Regular industry conference attendance
  4. An LMS platform

11. Your personal operating rhythm for staying current should optimise for…

  1. Reading through absolutely everything
  2. Filtered, recurring, bounded signal
  3. Constantly following the influencers
  4. Daily tool-switching

12. Pitching the iteration model to a real stakeholder (this week's mission) matters because…

  1. Practice makes perfect
  2. Real objections reveal actual constraints
  3. The stakeholders genuinely enjoy the pitches
  4. It completes the rubric

Module 18 — Capstone: Build, Govern, and Tell the Story

Outcome: Synthesise everything: identify, design, build, and govern a real agentic AI solution — then present it like a leader.

Leadership lens: The capstone is your proof-of-leadership: a real problem, a working (or credibly prototyped) agentic solution, a governance posture, measured impact, and a story that lands with executives and engineers alike.

Apply-at-work mission — Ship the capstone: Complete your capstone in the Capstone Tracker: real problem, built solution or detailed prototype, governance + risk register, impact measurement, published portfolio, and a 10–15 min recorded walkthrough.

Reflection: Final entry: read your Week 1 reflection, then write to your past self. What did you most misunderstand about leading with AI 18 weeks ago?

Resources

Project — Capstone: Build, Govern, and Tell the Story (Capstone Milestone 6)

Identify a meaningful real problem in your work or community. Design and build a complete agentic AI solution (or detailed prototype) addressing it — including governance, monitoring, and scaling considerations from your weeks 6, 8, and 16 artifacts. Create a professional portfolio (site or Notion) showcasing the capstone, 3–4 other programme projects, and your leadership philosophy. Record a 10–15 minute video walkthrough. 🎯 This completes Capstone Milestone 6 — and the programme.

Deliverable: portfolio/w18-capstone.md + published portfolio link + video — the complete story: problem, solution, governance, impact, journey.

Assessment rubric

Criterion Weight What good looks like
Problem significance & fit 20% A real problem someone actually has, matched honestly to agentic capability — not a solution seeking a problem.
Solution completeness 30% Working system or credible detailed prototype; components, autonomy tiers, and human gates all deliberate and explained.
Governance & impact 25% Risk register, monitoring, and scaling plan attached; impact measured or honestly estimated with a method.
Portfolio & narrative 25% Published portfolio a stranger could assess; the video tells a leadership story — problem, judgment calls, results — not a feature tour.

Knowledge check (12 questions)

Self-test prompts. Answers and explanations are not published here — take the quiz at https://ragentic.netlify.app/#/courses/agentic-ai-leadership to check yourself.

1. The strongest capstone problem selection criterion is…

  1. Sheer technical impressiveness
  2. A real, well-fitting pain
  3. Trending new technology use
  4. Maximum possible agent count

2. A capstone with brilliant engineering but NO governance section signals…

  1. Efficient prioritisation
  2. Incomplete leadership: capability without accountability is exactly what this programme argues against
  3. Trust in the model
  4. Agile thinking

3. Measuring capstone impact honestly, when full data isn't available, means…

  1. Simply skipping impact claims
  2. A method and honest estimate
  3. Using the biggest defensible number
  4. Quoting from various analyst reports

4. Your portfolio's primary audience design should target…

  1. Various other programme graduates just like you
  2. A sceptical decision-maker assessing you fast
  3. Search engines
  4. AI researchers

5. The video walkthrough should be structured as…

  1. A feature-by-feature screen tour
  2. A leadership narrative
  3. A full tutorial for rebuilding it
  4. A slide deck reading

6. "Leadership story" framing of your 18 weeks means…

  1. Simply listing all the completed modules
  2. Articulating how your thinking changed
  3. A plain chronological diary of the whole thing
  4. Certificates displayed

7. Including FAILURES in the portfolio (what didn't work) is…

  1. Career suicide
  2. A credibility multiplier: honest failure analysis
  3. It is really just some entirely optional padding for it
  4. Only for engineers

8. The multi-agent SDR system, ops command center, and governance dashboard capstone directions share…

  1. They all happen to use the very same tools
  2. End-to-end scope including oversight
  3. A strong shared sales-department focus
  4. Low difficulty

9. Presenting to executives vs engineers differs in that executives primarily need…

  1. A great many more acronyms used
  2. Outcomes, risks, decisions
  3. Architecture diagrams first
  4. Much longer meeting sessions

10. The "letter to your Week-1 self" reflection exercise exists because…

  1. Sentimentality
  2. Articulating your own change consolidates it — and reveals the misconceptions you'll now recognise in others
  3. It fills the last week
  4. Tradition

11. After the programme, the highest-value habit to keep is…

  1. Weekly quiz-taking
  2. The build-reflect-apply loop: hands-on testing of new capabilities, journaled reflection, and workplace application
  3. Reading more newsletters
  4. Collecting certificates

12. Sharing your capstone publicly ("build in public") primarily buys you…

  1. Likes
  2. Compounding credibility and a network
  3. A definite risk of plagiarism by others
  4. Nothing measurable

Toolkits

Agentic Workflow Canvas

Unlocks in module 6.

Map any process for agents: inputs, decisions, tools, exceptions, and the human gates that stay.

# Agentic Workflow Canvas

**Process name:**   **Owner:**   **Date:**

## 1. Outcome
What does "done, and done well" look like? Who consumes the output?

## 2. Trigger & inputs
What starts the process? List every input, its source, and its format (including the messy ones).

## 3. Steps & decisions
| # | Step | Actor (human/agent) | Decision criteria (if a decision) | Tools/data used |
|---|------|--------------------|-----------------------------------|-----------------|
| 1 | | | | |

## 4. Exception paths
For every step: what can go wrong, how is it detected, and where does it route (retry / fallback / escalate)?

| Step | Failure mode | Detection | Route |
|------|-------------|-----------|-------|

## 5. Human judgment gates (minimum 5 candidates, keep the real ones)
| Gate | Why human? (stakes / ambiguity / irreversibility) | What the human sees to decide |
|------|---------------------------------------------------|-------------------------------|

## 6. Observability
What is logged at each step? How would you reconstruct a run a week later?

## 7. Kill switch
How do you stop this workflow fast, and what is the manual fallback while it's stopped?

Pilot-to-Scale Maturity Assessment

Unlocks in module 8.

Grade any AI pilot on the 4-level model and plan the jump to the next level.

# Maturity Assessment — Experiment → Pilot → Operational → Scaled

**Workflow/pilot:**   **Assessed by:**   **Date:**

## Level definitions
1. **Experiment** — works on curated examples, builder-operated, no owner
2. **Pilot** — real users, real inputs, builder still on call, success criteria defined
3. **Operational** — named owner, error handling, integrated into real systems, support process exists
4. **Scaled** — multiple teams/contexts, portfolio governance, funded as a product

## Assessment
| Dimension | Evidence today | Level (1–4) |
|-----------|---------------|-------------|
| Ownership (who is accountable when it breaks?) | | |
| Input reality (curated vs whatever arrives) | | |
| Error handling & exceptions | | |
| Integration (real systems vs copy-paste) | | |
| Monitoring & metrics | | |
| Funding model (project vs product) | | |

**Overall current level:**   **Target level (6 months):**

## Blockers to next level (be specific)
1.
2.
3.

## Redesigned SOP (attach)
What does the agent do, what do humans do, how do exceptions route?

## Human+agent KPIs (efficiency AND quality)
| KPI | Type | Measurement method | Target |
|-----|------|--------------------|--------|

Governance Playbook & Risk Register

Unlocks in module 16.

Autonomy tiers, monitoring signals, escalation, rollback — plus the risk register for any agentic system.

# Agentic Governance Playbook & Risk Register

**System:**   **Owner:**   **Steward:**   **Date:**

## Autonomy tiers
| Tier | Definition | Examples for this system |
|------|-----------|--------------------------|
| Free | Agent acts, logs, no approval | |
| Gated | Agent prepares, human releases | |
| Forbidden | Agent may never do this | |

## Monitoring signals
| Signal | Threshold | Who is alerted | First response |
|--------|-----------|----------------|----------------|
| Error / correction rate | | | |
| Escalation rate | | | |
| Cost per run / day | | | |
| Behaviour drift vs baseline | | | |

## Escalation path
Who gets paged, in what order, with what authority?

## Rollback procedure (must work at 2am without the builder)
1. Kill switch location + who may pull it:
2. Manual fallback process while stopped:
3. Communication plan (users, stakeholders):

## Risk register
| # | Risk | Type (hallucination / action / compliance) | Likelihood | Impact | Owner | Control |
|---|------|--------------------------------------------|-----------|--------|-------|---------|
| 1 | | | | | | |
| 2 | | | | | | |
| 3 | | | | | | |

## Recourse
How does an affected person challenge an outcome, and which human can change it?

Cognitive Capital Audit

Unlocks in module 14.

Classify two weeks of your AI usage: thinking deepened (capital) or replaced (debt)?

# Cognitive Capital Audit

**Period reviewed:**   **Date:**

## Usage log
| Use of AI (significant instances) | What I did first, myself | Capital or Debt? | Evidence |
|-----------------------------------|--------------------------|------------------|----------|
| | | | |

**Capital** = my understanding/skill is stronger after this pattern of use.
**Debt** = I can no longer comfortably do (or judge) this without the tool.

## Debt areas (top 3)
| Skill degrading | Evidence | Do I want it back? | Repayment plan (practice) |
|-----------------|----------|--------------------|-----------------------------|

## My Thinking Amplification Protocol
Trigger (which decisions/documents):
1. Draft my own position first (timebox:   min)
2. Ask AI to attack it: weakest assumptions, missing perspectives, disconfirming evidence
3. Revise and record what changed
Practice commitment:   days/week for   weeks

## Team norms I will model
- Think-then-ask on:
- "The AI said so" is never a complete justification for:

Orchestrator–Architect–Steward Role Map

Unlocks in module 13.

Map your team against the three roles that scale; find the gap; draft the missing job.

# Role Mapping — Orchestrator / Architect / Steward

**Team:**   **Mapped by:**   **Date:**

## Role definitions
- **Orchestrator** — aligns humans + agents across silos; owns handoffs, context, and cadence
- **Architect** — designs workflows, decision logic, and system structure
- **Steward** — owns trust, safety, accountability, and governance of human+agent systems

## Current coverage
| Person / function | Orchestrator | Architect | Steward | Evidence |
|-------------------|--------------|-----------|---------|----------|
| | | | | |

## Failure symptoms observed (map to missing role)
| Symptom | Points to gap in |
|---------|------------------|
| Governance incidents despite fast shipping | Steward |
| Fragile ad-hoc undocumented workflows | Architect |
| Silos, duplicated agent work, broken handoffs | Orchestrator |
| (your observations) | |

## My gravitational role
Under pressure I default to:   Evidence:
Development plan for my weakest role:

## Draft job description for the missing role
Title:
Mission:
Responsibilities (5):
Success at 6 months / 12 months:

Continuous Transformation Playbook

Unlocks in module 17.

Operating rhythm, funding model, and kill criteria for a persistent AI iteration capability.

# Continuous Transformation Playbook

**Scope (team/org):**   **Author:**   **Date:**

## Operating rhythm
| Cadence | Activity | Output |
|---------|----------|--------|
| Weekly | Frontier scan + one hands-on test | Tested-capability note |
| Monthly | Evaluate 1–2 capabilities against live workflows | Adopt / watch / reject decision |
| Quarterly | Portfolio review: retire, scale, re-govern | Updated workflow portfolio |

## Iteration capability
Team shape (roles, % time):
Budget envelope:
Decision rights (what they may change without approval):

## Kill criteria (pre-committed)
An experiment stops when:   (e.g. no measurable value after N cycles, cost/quality regression, unowned risk)

## Change-fatigue guards
Stable anchors that do NOT change (purpose, quality bar, core rituals):

## Scenarios 2028–2030 (three futures)
| Scenario | What the industry looks like | What we start doing NOW |
|----------|------------------------------|--------------------------|
| | | |

Apply-at-Work Mission Log

Unlocks in module 1.

The running record of every weekly mission: what you did at work, what happened, what you learned.

# Apply-at-Work Mission Log

| Week | Mission | What I actually did | Outcome / reaction | What I'd do differently |
|------|---------|---------------------|--------------------|-------------------------|
| 1 | Hire your first AI teammate | | | |
| 2 | Brief your team on the inflection | | | |
| 3 | Automate one monitoring task | | | |
| 4 | Ship a tool your team uses | | | |
| 5 | Connect a prototype to a real system | | | |
| 6 | Canvas a real workflow with a colleague | | | |
| 7 | Stand up a multi-agent system | | | |
| 8 | Maturity-assess a real pilot | | | |
| 9 | Build a lead/ops research agent | | | |
| 10 | Repurpose content with brand guardrails | | | |
| 11 | Design a review-gated HR/Finance workflow | | | |
| 12 | Run a WHY–WHAT–HOW conversation | | | |
| 13 | Role-map your team | | | |
| 14 | Audit your AI usage | | | |
| 15 | Run an adoption audit | | | |
| 16 | Draft a real risk register | | | |
| 17 | Pitch an iteration team | | | |
| 18 | Ship the capstone | | | |

Capstone

Enterprise · Multi-agent SDR system

Research + qualification + personalised outreach preparation, with compliance gates and a human release step.

Enterprise · Intelligent operations command center

Monitors multiple data sources, correlates signals, and triggers coordinated (gated) responses.

Enterprise · Customer success intervention system

Detects churn-risk signals and orchestrates intervention workflows with human owners.

Leadership & Governance · Agentic governance framework + monitoring dashboard

A real organisation's autonomy tiers, risk register, and a live monitoring view over its agent workflows.

Leadership & Governance · Role-based agent team design for a department

Orchestrator + Architect + Steward blueprint applied to one department, with workflows and decision rights.

Leadership & Governance · Cognitive capital development programme

A team-level programme using AI as thinking partner (not answer engine): protocols, norms, and measurement.

Your own · Your own real problem

The best capstone: a meaningful problem from your work or community that an agentic approach genuinely fits.

Milestones

Portfolio checklist