The Topic Catalog · 6

AI for Professional & Knowledge Workers

Last updated · 9 min read

6.1 The Maturity Ladder: Chat → Assistants → Connected Assistants → Workflows → Agents → Autonomous Business Processes

Priority: Must Understand

Executive Definition: Six rungs describe how AI moves from novelty to operational infrastructure: Chat (ad hoc prompting, no system access), Assistants (embedded in one app: Word, Excel), Connected Assistants (cross-system retrieval with permissions, e.g., enterprise search), Workflows (deterministic multi-step processes with an AI step inserted), Agents (goal-directed, tool-using, makes some decisions independently), and Autonomous Business Processes (agents run an end-to-end process with human oversight by exception, not by step). Most enterprise deployments today sit at rungs one through three; vendor marketing routinely implies rungs five and six.

Why It Matters: Gartner calls the gap between marketed and actual autonomy "agentwashing" (most products sold as "agents" are still reactive assistants) and separately predicts over 40% of agentic AI projects will be canceled by end of 2027 due to unclear ROI, escalating cost, and weak risk controls (Gartner). MIT's Networked Agents and Decentralized AI (NANDA) initiative found 95% of enterprise GenAI pilots produced no measurable return, and traced the divide not to model quality but to whether organizations redesigned the workflow around the tool or just added a chat box on top of an unchanged process (MIT NANDA). This ladder is the vocabulary for scoping a project honestly and refusing inflated claims.

What I Need to Understand:

  • Each rung requires different governance: approval-per-action at "Workflows," approval-by-exception at "Autonomous Business Processes": the risk model changes at every step, not just the tech.
  • Moving up a rung is a process-redesign exercise, not a model swap or a licensing upgrade.
  • "Adding chat" to an existing workflow is a rung-one or rung-two change even when marketed as agentic.
  • McKinsey found high performers were three times more likely to have "fundamentally redesigned workflows" than typical adopters (74% vs. 25%), and that this (not tool adoption) is what correlates with financial impact (McKinsey).
  • Vendor claims of "agent" should be interrogated against Gartner's stage definitions before budget commitment.

Questions I Should Be Able to Ask My Team:

  1. Which rung is each of our current AI deployments actually operating at, and who verified that classification independent of the vendor's own description?
  2. For anything sold to us as "agentic," what decision is the system actually making autonomously, versus what is still a scripted step with an LLM call inserted?
  3. What changed in the underlying business process (not just the tooling) for our highest-value AI deployment, and can we show a before/after process map?

Technologies / Standards / Companies to Know: Gartner Hype Cycle for Agentic AI, MIT NANDA "State of AI in Business," Microsoft Agent 365, Forrester Adaptive Process Orchestration.

Recommended Learning:

Time Investment: 2-3 hours


Priority: Must Understand

Executive Definition: Two distinct product categories get conflated under "AI assistant." App-embedded copilots (Microsoft 365 Copilot in Word/Excel, Gemini in Docs/Sheets) work inside one application. Enterprise search/assistant platforms (Glean, Google Gemini Enterprise, Copilot with Microsoft Graph connectors, Claude Enterprise with connectors) index content across many systems and answer questions grounded in whatever the requesting user is already permitted to see. The second category's value and risk both hinge on one mechanism: permission-aware retrieval: not model quality.

Why It Matters: Glean's documentation is explicit that answers are only as good as the underlying permissions model: the system builds a knowledge graph from connector-crawled content, activity, and identity data, and filters every result through existing access controls before it reaches the model (Glean). Get this wrong and you either expose data across permission boundaries or produce thin, unhelpful answers because whole systems weren't indexed: both are governance failures dressed up as AI-quality complaints.

What I Need to Understand:

  • App-embedded copilots and cross-system enterprise search solve different problems and are frequently bought and evaluated as if interchangeable.
  • Connector/indexing coverage (which systems are actually crawled) determines usefulness far more than which model powers the assistant.
  • Permission propagation lag (offboarding, role changes, revoked access) is a real security gap if the AI's index doesn't sync in near-real time with source-system ACLs (Glean).
  • "Hallucination" complaints in an enterprise-search context are often retrieval/grounding failures (wrong or missing source document), not model failures: the fix differs accordingly.
  • Licensing is typically per-seat and stacks on top of, not instead of, existing SaaS licenses: check for overlap before buying a second search layer.

Questions I Should Be Able to Ask My Team:

  1. What percentage of our critical knowledge systems are actually indexed by this assistant, and what's explicitly excluded?
  2. When someone's access is revoked or changes, how long until that's reflected in what the AI can retrieve and answer from?
  3. Can we trace any given answer back to the specific documents and the permission check that allowed them to be used?

Technologies / Standards / Companies to Know: Microsoft 365 Copilot + Graph connectors, Google Gemini Enterprise (formerly Agentspace), Glean, Claude Enterprise + Anthropic connectors/MCP.

Recommended Learning:

Time Investment: 2-3 hours


6.3 Deep Research Agents & Document/Presentation/Spreadsheet Generation

Priority: Monitor

Executive Definition: Deep research agents (OpenAI Deep Research, Google Gemini Deep Research, NotebookLM Deep Research, Microsoft Researcher in Copilot Notebooks) run asynchronously for minutes rather than seconds: they plan a research strategy, search and read across many sources, and return a cited report: closer to delegating to a junior analyst than to prompting a chatbot. These increasingly connect directly to document generation, turning that output into a Word doc, slide deck, or spreadsheet.

Why It Matters: OpenAI's own account of Deep Research describes it finding, analyzing, and synthesizing "hundreds of online sources" into an analyst-grade report in 5-30 minutes, built on a reasoning model tuned specifically for multi-step browsing tasks (OpenAI). The genuine capability shift is real, but benchmark scores (GAIA, Humanity's Last Exam) measure general research aptitude, not accuracy on your specific domain, and every generated report still needs a human citation-verification pass before external use.

What I Need to Understand:

  • The asynchronous, multi-minute execution model changes how this fits into a workflow: it's not a replacement for instant chat, it's a delegated task.
  • Grounding source matters: public-web research (OpenAI, most consumer deep research) versus enterprise-data-grounded research (Gemini Deep Research pulling from indexed internal sources) are different risk and value propositions (Google Cloud).
  • Citations must be independently verified before a report leaves the building: misattribution and fabricated sources remain a known failure mode of these systems.
  • This is distinct from templated document automation (mail merge, standard reporting), which is deterministic, cheaper, and should not be replaced by an agent for the sake of it.
  • "High steerability" (custom tone, structure, format) is a real differentiator between vendors and worth testing on your own use cases, not vendor demos.

Questions I Should Be Able to Ask My Team:

  1. When this tool cites a source, what's our process for verifying that citation before the output is used in an external-facing document?
  2. Is this agent grounded in public web data, our own enterprise data, or both, and does that vary by user or by query?
  3. What's the measured time and cost per report on our own use cases, compared to an analyst doing the same task: not the vendor's benchmark numbers?

Technologies / Standards / Companies to Know: OpenAI Deep Research, Google Gemini Deep Research / NotebookLM, Microsoft Researcher (Copilot Notebooks), Anthropic Claude web search and citations.

Recommended Learning:

Time Investment: 1 hour


6.4 Meeting Intelligence, Departmental Digital Workers & Function-Specific Agents (HR, Finance, Legal, Procurement, Sales, Operations, Program Management)

Priority: Should Understand

Executive Definition: Two related but distinct categories. Meeting intelligence tools (Otter, Microsoft Copilot in Teams, Google Gemini in Meet, Read.ai, Fireflies) transcribe, summarize, and extract action items from calls. Departmental "digital worker" agents are embedded directly in a system of record and scoped to its data: Salesforce Agentforce (sales/service), SAP Joule (finance, HR, supply chain), Workday agents (HR), ServiceNow agents (IT/operations). Because they're bounded to one platform's data model, these are typically the most concretely deployed and measurable category of enterprise AI today.

Why It Matters: SAP has shipped roughly 15 named Joule agents across finance, HR, and supply chain functions (Techzine; independent commentary from HR analyst Josh Bersin covers the same rollout: Bersin), and McKinsey found large enterprises (>$1B revenue) scaling AI agents at 40%, versus 22% flat for smaller organizations: the gap that matters for a mid-size enterprise sizing its own ambitions realistically (McKinsey). The recurring trap is the same as elsewhere: a meeting summarizer or departmental agent that reproduces a report you already generate is not transformation: it only counts if the underlying process changed.

What I Need to Understand:

  • Meeting intelligence tools introduce their own governance surface: recording consent, retention policy, and who can query historical meeting transcripts across the org.
  • Departmental agents are generally vendor-locked to that platform's data model: cross-platform orchestration (e.g., a Workday agent triggering a SAP process) is immature and should not be assumed.
  • Pricing for departmental agents is increasingly per-agent or per-outcome/conversation rather than per-seat, which is harder to forecast and budget than traditional SaaS licensing.
  • Function heads (HR, Legal, Procurement) will be sold these agents directly by vendors, often bypassing central IT/architecture review: a coordination point is needed before contracts land.
  • Maturity varies sharply by function: sales/service (Agentforce) and IT operations (ServiceNow) are furthest along; legal and procurement agent tooling is comparatively immature and less independently validated.

Questions I Should Be Able to Ask My Team:

  1. Who owns governance for meeting recordings and transcripts: retention period, consent, and who can search past meetings org-wide?
  2. For each departmental agent under evaluation, what specific manual process does it replace, and can we see the actual before/after workflow, not just the vendor demo?
  3. How does this agent's pricing scale if adoption succeeds and usage triples: is that cost modeled anywhere yet?

Technologies / Standards / Companies to Know: Microsoft Copilot in Teams, Google Gemini in Meet, Otter.ai, Read.ai, Salesforce Agentforce, SAP Joule, Workday agents, ServiceNow AI agents.

Recommended Learning:

Time Investment: 1 hour