The Topic Catalog · 21
Emerging Technology Radar
For each item below: classification (Near-term / Developing / Highly Uncertain-Speculative), what it is, why it matters enough to watch, and the concrete signal that would indicate it's moving from speculative toward real.
21.1 Autonomous Software Development (issue-to-PR-to-deploy)
Classification: Developing What it is: Coding agents (Devin, GitHub Copilot coding agent, Cursor/Cognition background agents, Claude Code, etc.) that take a ticket, write and test code, open a PR, and in some pipelines merge and deploy with little or no human step in between. Why watch it: If reliable, this changes engineering headcount math and where review/QA effort goes, but reliability, not model capability, is the binding constraint at present. Signal to track: SWE-bench Verified scores and, more importantly, real production incident/rollback rates disclosed by early adopters running agents with merge authority (not sandboxed benchmarks).
21.2 Self-Healing Software
Classification: Developing What it is: AI systems that detect a production fault, diagnose root cause, and remediate (rollback, restart, patch, reconfigure) without a human in the loop. Why watch it: Narrow, bounded self-healing (auto-rollback, known-failure runbooks) is already viable and cuts MTTR; general-purpose autonomous remediation of novel failures is not: the gap between vendor claims and audited outcomes is wide. Signal to track: Whether a major APM/AIOps vendor (Datadog, PagerDuty, Dynatrace) publishes audited MTTR/false-remediation data for closed-loop (no human approval) fixes, not just "AI-assisted" ones.
21.3 Software Factories (fleets of coding agents as a production system)
Classification: Developing What it is: Treating dozens-to-hundreds of coding agents as a managed production system (with queues, review gates, evals, and observability) rather than one agent per developer. Why watch it: This is the operating-model question behind autonomous coding: it determines whether agent output scales safely or just scales incident volume. Signal to track: Vendors like Factory.ai publishing fleet-scale reliability/throughput metrics, and whether a large enterprise (not a startup) discloses agent-fleet headcount-equivalent numbers.
21.4 Agent Marketplaces
Classification: Developing What it is: Storefronts for pre-built third-party agents that plug into an enterprise's stack: Salesforce's AgentExchange (successor to AppExchange) is the clearest concrete example, alongside Microsoft's and Google's agent catalogs. Why watch it: Marketplaces shift agent risk from "we built it" to "we procured it," which raises new vendor-vetting, security, and liability questions procurement isn't yet set up for. Signal to track: Whether AgentExchange-style marketplaces publish real usage/trust metrics (installs, incident disclosures) rather than just listing counts. (salesforce.com)
21.5 Agent Discovery Protocols
Classification: Near-term What it is: Standards (Anthropic's Model Context Protocol for tool/data access; Google's Agent2Agent/A2A for agent-to-agent interop, now under the Linux Foundation) that let agents find and call each other's capabilities. Why watch it: MCP has become the de facto integration layer for agent tooling in about two years: protocol consolidation (or fragmentation) directly affects integration cost and vendor lock-in risk. Signal to track: The Linux Foundation reports A2A passed 150+ member organizations with early enterprise production use; watch whether MCP and A2A converge, merge governance, or start competing for the same layer. (linuxfoundation.org, developers.googleblog.com)
21.6 Agent Reputation Systems
Classification: Highly Uncertain/Speculative What it is: Mechanisms to score an agent's trustworthiness/track record before letting it transact or act on your behalf: proposals range from centralized marketplace ratings to on-chain registries (e.g., Ethereum's ERC-8004 "trustless agents" proposal). Why watch it: Without reputation, agent-to-agent commerce and open marketplaces can't scale past known counterparties: it's a prerequisite, not a nice-to-have, for items 4, 7, 27 and 28. Signal to track: Any agent marketplace (Salesforce, Google, OpenAI) shipping a real reputation/score field that affects transaction routing, versus reputation remaining a marketing claim.
21.7 Agent Payments / Agentic Commerce
Classification: Developing What it is: Protocols letting an AI agent initiate and complete a purchase on a human's or business's behalf: Google's Agent Payments Protocol (AP2, with Mastercard, Coinbase, and others), and the OpenAI/Stripe Agentic Commerce Protocol behind ChatGPT's Instant Checkout. Why watch it: Real card networks and payment processors are already building rails for this; enterprise exposure (procurement agents making purchases) will arrive faster than most CFOs expect. Signal to track: Whether AP2/ACP transaction volume moves beyond consumer shopping demos into Business-to-Business (B2B) procurement, and whether card networks publish agent-initiated fraud/chargeback rates. (cloud.google.com, stripe.com)
21.8 Agent Identity Standards (beyond OAuth)
Classification: Developing What it is: Purpose-built identity systems for non-human actors: Microsoft's Entra Agent ID (reached GA in 2026) is the clearest example, giving each agent its own directory identity, lifecycle, and permissions distinct from a service account or a human's delegated OAuth token. Why watch it: Today's OAuth-based patterns weren't designed for entities that spawn sub-agents and act continuously; identity sprawl is already a named problem in early Entra Agent ID deployments. Signal to track: Cross-vendor interop: whether an Entra Agent ID can be recognized/verified by a non-Microsoft agent platform, not just within Microsoft's stack. (learn.microsoft.com)
21.9 Persistent Digital Workers (named, long-lived agent "employees")
Classification: Developing What it is: Agents given a persistent identity, role, and memory across sessions (framed explicitly as "digital labor" by Salesforce and similarly by Microsoft's Copilot agent framing) rather than stateless per-task tools. Why watch it: This is the framing vendors are using to sell per-seat/per-outcome pricing for agents; the actual persistence and memory quality (versus a rebranded chatbot) determines whether it's substance or packaging. Signal to track: Independent (non-vendor) case studies quantifying a named digital worker's task success rate and cost-per-outcome over months, not launch-day demos. (salesforce.com)
21.10 Autonomous Departments
Classification: Highly Uncertain/Speculative What it is: An entire business function (e.g., collections, tier-1 support, basic procurement) run predominantly by coordinated agents with a human only in an oversight role. Why watch it: This is the aggregation point where individual agent ROI either compounds into department-level headcount change or stays a productivity veneer: worth tracking because it's the scenario most consultancies (Deloitte, BCG) are modeling toward 2027-2028, not one broadly achieved today. Signal to track: A named, audited case of a function running with a materially reduced human headcount for over a year (not a pilot announcement). (deloitte.com)
21.11 Synthetic Organizations
Classification: Highly Uncertain/Speculative What it is: The theoretical end state of the above: an organization whose operating structure is substantially composed of interacting agents rather than human teams, sometimes discussed in academic/futurist contexts as multi-agent "societies." Why watch it: Worth naming precisely so it isn't confused with today's real (and much more limited) digital-worker and autonomous-department efforts: the absence of credible reporting on this happening anywhere in 2026 is itself the calibrating signal. Signal to track: Any peer-reviewed or analyst-verified example, as distinct from vendor thought-leadership content: none currently meets that bar.
21.12 AI-Native Enterprise Resource Planning (ERP)
Classification: Developing What it is: ERP vendors (SAP with Joule agents, Oracle Fusion, Microsoft Dynamics) adding agentic layers on top of existing systems of record, short of a ground-up re-architecture around agents. Why watch it: Most "AI-native ERP" today is agent bolt-ons to legacy data models: the strategic question is whether incumbents or challengers get to the re-architecture first, since that's where lock-in shifts. Signal to track: Whether SAP or Oracle ships agent-initiated transactions (e.g., autonomous PO creation/approval) as a default workflow rather than an assisted suggestion, with disclosed error rates.
21.13 AI-Native CRM
Classification: Developing What it is: Salesforce Agentforce and HubSpot Breeze are the two clearest branded pushes toward CRM where agents (not just automations) handle lead qualification, service cases, and follow-up autonomously. Why watch it: CRM is the highest-volume, lowest-stakes-per-transaction place to pressure-test agent autonomy in a live revenue system before extending it to higher-stakes domains. Signal to track: Independent (Gartner/Forrester) satisfaction and containment-rate data for Agentforce/Breeze deployments, versus vendor-published case studies.
21.14 Generative UI (interfaces generated on the fly)
Classification: Developing What it is: Interfaces assembled dynamically per query/task by a model rather than pulled from a fixed set of pre-built screens: Google Research has published this explicitly as a direction ("a rich, custom, visual interactive experience for any prompt"). Why watch it: If it works, it collapses the cost of building bespoke internal tools/dashboards; if it doesn't, it produces inconsistent, unreviewable UI sprawl: the design and accessibility tooling to govern this at scale doesn't exist yet. Signal to track: Whether a major consumer product (Search, Gemini app) ships generative UI as a default experience rather than an experimental toggle. (research.google)
21.15 Software Generated on Demand (ephemeral, single-use apps)
Classification: Developing What it is: Disposable, single-purpose applications a model generates for one task and discards: already visible in miniature via features like Claude's Artifacts or ChatGPT's Canvas generating a one-off tool mid-conversation. Why watch it: This challenges the economics of building and maintaining internal software at all for narrow, recurring-but-low-volume needs, but raises real questions about security review, data governance, and audit trail for code no one ever formally ships. Signal to track: Whether any enterprise formally sanctions (with a governance policy) employee-generated ephemeral apps touching real business data, versus treating all such use as shadow IT.
21.16 Natural-Language Programming (non-developers directing agents)
Classification: Developing What it is: "Vibe coding" (a term popularized by Andrej Karpathy in early 2025) where someone without formal engineering training describes intent in natural language and an agent (Replit, Lovable, Cursor, Claude Code, base44, etc.) produces working software. Why watch it: This is the most direct threat/opportunity to the "who is allowed to build software" boundary in a large engineering org: shadow engineering by knowledge workers is already happening, whether or not IT has a policy for it. Signal to track: Whether your own knowledge workers are already producing production-adjacent tools this way (check for unsanctioned Replit/Lovable/Claude-built internal tools) before deciding this is someone else's problem.
21.17 Edge AI
Classification: Near-term What it is: Running inference on devices/gateways at the network edge rather than in the cloud, for latency, cost, connectivity, or data-residency reasons. Why watch it: This is a mature, well-funded infrastructure category already embedded in manufacturing, retail, and industrial IoT deployments: the executive question is procurement/architecture, not "should we watch this." Signal to track: Total cost of ownership crossover points as edge hardware (NPUs) gets cheaper: track your own inference cost mix moving from cloud API calls to edge/on-prem for high-volume, low-complexity tasks.
21.18 Local / On-Device Models
Classification: Near-term What it is: Capable small models running entirely on a phone, laptop, or PC without a network call: Apple Intelligence (reportedly now built in part on licensed Google Gemini models per 2026 reporting), Gemini Nano, Microsoft's Phi family, and Llama variants are the concrete examples. Why watch it: This determines the floor for what "AI" means when connectivity, latency, or data-privacy requirements rule out a cloud call: increasingly relevant for regulated data and field/offline use cases. Signal to track: Whether your device fleet (laptops issued to your knowledge workers) ships with a capable local model by default, changing what needs cloud API budget at all. (appleinsider.com)
21.19 Embodied AI
Classification: Highly Uncertain/Speculative What it is: AI systems that perceive and act in the physical world through a body (robot, drone, vehicle) rather than through text/screen interfaces, using the same foundation-model techniques as LLMs. Why watch it: Almost entirely irrelevant to a software/knowledge-work enterprise directly, but relevant as a leading indicator of foundation-model capability generalizing beyond language, and China's state-level push (per Merics reporting) is a geopolitical signal worth tracking even if you never buy a robot. Signal to track: Independent, non-demo-reel evidence of embodied agents completing multi-step real-world tasks reliably outside controlled lab/showroom settings. (merics.org)
21.20 Robotics (AI-driven, foundation-model-based)
Classification: Developing What it is: Vision-language-action (VLA) foundation models controlling physical robots: Figure AI's Helix model and Physical Intelligence's π-0 are the clearest named examples, alongside Chinese firms like AgiBot. Why watch it: Same reasoning as embodied AI generally, but this is the layer with actual funded companies, named models, and pilot deployments (logistics, manufacturing) rather than pure research: worth distinguishing from the broader speculative category above. Signal to track: A named humanoid/robot deployment moving from a single-site pilot to multi-site commercial contract with disclosed unit economics. (en.wikipedia.org)
21.21 World Models
Classification: Highly Uncertain/Speculative What it is: Models trained to build an internal predictive simulation of how the physical/visual world evolves (rather than just predicting text), exemplified by DeepMind's Genie line and Fei-Fei Li's World Labs (Marble). Why watch it: Proponents argue world models are a precondition for robust embodied AI and long-horizon planning; skeptics note none has yet demonstrated enterprise utility beyond generating interactive video/game environments. Signal to track: Any world model demonstrating a real economic use case beyond content/simulation generation (e.g., materially improving robotic task success or supply-chain simulation) versus remaining a research demo.
21.22 Digital Twins (AI-enhanced)
Classification: Developing What it is: Physics- or data-based simulations of a real asset, process, or system (a factory line, a supply chain, a building), now increasingly layered with AI for prediction and what-if reasoning: Nvidia Omniverse and Siemens Xcelerator are the established commercial platforms. Why watch it: Unlike most items on this list, digital twins are already operationally mature in manufacturing/industrial contexts: the "watch" item is specifically the AI layer (generative simulation, agent-driven scenario testing) being added on top of an already-proven category. Signal to track: Whether AI-driven scenario generation in a digital twin platform demonstrably shortens a real capital-planning or process-redesign cycle, versus remaining a visualization upgrade.
21.23 Synthetic Data (for training/eval at enterprise scale)
Classification: Near-term What it is: Artificially generated data used to train or evaluate models when real data is scarce, sensitive, or imbalanced: already central to frontier model training (Nvidia's Nemotron pipelines, Microsoft's Phi models trained substantially on synthetic data). Why watch it: For a large enterprise engineering organization, synthetic data is the practical near-term path to fine-tuning/evaluating internal models on sensitive data (code, customer records) without exposing the real thing: governance and quality-control practices matter more than the underlying tech. Signal to track: Whether your own model-eval or fine-tuning pipelines have an explicit synthetic-data quality/bias-audit step, not just a generation step.
21.24 AI for Scientific Discovery
Classification: Developing What it is: AI systems directly accelerating scientific research: DeepMind's AlphaFold (2024 Nobel Prize in Chemistry to Hassabis and Jumper), GNoME for materials discovery, and Google's "AI co-scientist" system (announced 2025) are the clearest proof points. Why watch it: These are the most credible, least-hyped evidence that foundation models generalize beyond language to genuine scientific value: relevant to any enterprise with an R&D, pharma, materials, or advanced-engineering function. Signal to track: Whether AI-assisted discoveries (new materials, drug candidates) move from computational prediction into validated real-world results at a rate distinguishable from normal R&D baseline.
21.25 Real-Time Multimodal Agents (voice + vision + action, low latency)
Classification: Developing
What it is: Agents that see, hear, and respond in real time with low enough latency for natural conversation and action: OpenAI's GPT-4o Realtime API and Google's Project Astra/Gemini Live are the current reference implementations.
Why watch it: This is the interface layer that could replace app/dashboard-based interaction for many knowledge-work tasks (a live voice+screen assistant), but reliability and cost at scale are still unproven outside demos.
Signal to track: Adoption of these APIs inside real enterprise customer-service or field-service workflows (not consumer demos), with disclosed latency and error rates.
21.26 Autonomous Cybersecurity (AI-driven defense/response)
Classification: Developing What it is: AI agents that detect, triage, and in some cases autonomously respond to security threats: Google's "Big Sleep" agent found a real-world exploited vulnerability in 2024, and Microsoft Security Copilot and CrowdStrike Charlotte AI now offer agentic response capabilities. Why watch it: Security is one of the few domains where both attacker and defender are racing to adopt agentic AI simultaneously: falling behind here has asymmetric downside compared to most other items on this list. Signal to track: Whether your Security Operations Center (SOC)/EDR vendor's "autonomous response" features are enabled with real remediation authority in your environment, versus running in detect-and-recommend mode only.
21.27 Machine Customers (Gartner's term)
Classification: Developing What it is: AI agents that shop, negotiate, and buy on behalf of a human or organization: Gartner projected in 2023 that 20% of inbound customer-service contact volume would come from machine customers by 2026. Why watch it: If even partially realized, this changes customer-service design (agents don't need the same UX as humans) and demand-forecasting assumptions: worth checking against Gartner's own actual-vs-predicted reporting now that 2026 has arrived. Signal to track: Gartner's updated (post-hoc) assessment of whether the 20% figure was reached, and whether your own contact-center logs show a rising share of API/agent-originated (vs. human) interactions. (gartner.com)
21.28 Agent-to-Agent Commerce
Classification: Highly Uncertain/Speculative What it is: Transactions initiated and completed between two AI agents (a buying agent and a selling agent) with no human approving the specific transaction: the combination of agent discovery (A2A/MCP), payments (AP2/ACP), and reputation systems above. Why watch it: This is the composite, furthest-out scenario on this list: each prerequisite protocol exists individually in 2026, but no credible reporting yet shows agent-initiated B2B or B2C transactions happening at meaningful volume without a human confirming the specific purchase. Signal to track: The first disclosed enterprise procurement transaction completed end-to-end by autonomous agents on both sides, verified by something more than a vendor press release.