National AI Challenge 2026 · Idea Hydration
CoWork
Enterprise AI Agent Workplace

AI is eating SaaS. CoWork is the harness that realizes it.

Enterprise SaaS was never valuable just because it was software. It was valuable because it encoded organizational problem-solving — roles, approvals, compliance, and coordination. "AI eating SaaS" means eating that layer, not the tech stack. CoWork re-embodies those organizational functions as a runtime for agents, so the work itself — not a tool — becomes the product.

Theme: AI & Agents — Strategy, Processes and Roles Stage 0 → 12 · Draft · not final judge copy Source: Idea Hydration Scaffold v3
Part 4 · One-line snapshot

Build-readiness in five lines

The scaffold requires these five answers to stay reducible to one sentence each — if they can't be, the stage isn't locked.

The Problem

Coordination, not cognition

Engineering work that spans more than one team, system, or compliance boundary is throttled by coordination and governance — and today's tools are shaped for a single developer, not that organizational work.

The Audience

Platform & security leads

They own the cross-cutting change, feel the coordination tax today, and are triggered to act the moment a breaking change or CVE can't wait on other teams' backlogs.

The Core Value

Steer once, scale everywhere

Dispatch a cross-boundary change, steer the exceptions with human judgment, and get a signed record of everything — shipping in days what today takes teams and months.

The Mechanics

Agent acts, human steers

Agents act autonomously inside a scoped, ephemeral sandbox; humans approve and steer at policy gates; brokers enforce identity, consent, and a tamper-evident audit trail.

The Constraints

Two weeks, offline-safe

Real: interactive demo + deterministic engine. Simulated: cloud sandboxes, live inference, Slack. Beachhead is enterprise platform/security — stated, with a path outward.

Parts 1–2 · The 13 hydration stages

Stage by stage

Each stage produced a Decision Record. "Locked" is the agreed position; "Open / uncertain" is what remains unresolved and must not be silently assumed.

0

Establish the Starting Point

What does the artefact actually say, before we improve it?
Locked
CoWork is an audited, ephemeral cloud workplace that decouples agent execution from local machines via six abstractions: ComputeRuntime, WorkspaceVolume, InteractionSurface, IdentityBroker, ConsentBroker, AuditLogger.

Original intent (preserved)

  • Problem hypothesis: enterprise engineering work is blocked at organizational seams, not at the code-writing step.
  • Audience: feature developers, platform engineering, security/compliance.
  • Value proposition (new framing): realizing "AI is eating SaaS" — where SaaS means the organizational problem-solving layer, not the tool.

Open / unknown at stage 0

  • Whether the six abstractions are a product or merely infrastructure plumbing.
  • Which single buyer feels the pain sharply enough to pay.
  • What "harness" concretely means to a judge in 30 seconds.
1

Hydrate the Problem

What is actually going wrong, and why does it persist?
Locked
The underlying problem is coordination, not cognition. Cross-boundary work — breaking changes, dependency migrations, security upgrades — crosses repos, teams, and governance boundaries, and no single tool is shaped to act across them.

What we agreed

  • Symptom: slow migrations, broken shared staging, months of backlog.
  • Cause: the work is organizational — who owns the boundary, who approves, who audits — not just technical.
  • Why it persists: SaaS decomposed into point tools, each solving one function but recreating coordination burden at the seams; the org layer was never made machine-executable.
  • Plain-language problem (no solution embedded): "Cross-boundary engineering work is throttled by coordination and governance, and today's tooling is built for one developer, not for that organizational work."

AI materially changes it?

  • Yes — but only with a harness. Agents can hold and act across many systems at once, which was impractical before; but ungoverned agents are just faster single developers.
  • open The economic cost needs a defensible anchor — we use the MuleSoft 39% figure below as a proxy, not a precise cost model.
Evidence. MuleSoft's 10th annual Connectivity Benchmark (2025) finds ~39% of IT team time goes to building and testing custom integrations — i.e. wiring systems together, not shipping features. Source: MuleSoft Connectivity Benchmark, 2025 · vendor survey.
2

Hydrate the Audience

Who feels it, and what is the exact trigger that makes them act?
Locked
Primary persona: the platform / security lead who owns a cross-cutting change. The trigger event is a breaking change or CVE with a hard deadline that can't be done in one repo and can't wait on other teams.

Who they are

  • User: platform/tech lead (Scenario 1's Alex, Scenario 2's David).
  • Beneficiary: the engineering org that stops breaking staging and waiting on backlogs.
  • Buyer / decision-maker: VP Engineering, Platform, or CTO; security/compliance gate the trust.
  • JTBD: "When I must change something across service boundaries, I want to dispatch and steer it with full visibility and a signed record, so I can ship without breaking others or waiting on their backlogs."

Awareness & objection

  • Awareness state: problem-aware, but solution-unaware and agent-skeptical — they will not trust ungoverned agents.
  • Primary objection: "How do I know what the agent did, and who is accountable?"
  • Adoption barrier: security sign-off; the audit trail is the thing that removes it.
3

Hydrate Context & Evidence

What is the defensible evidence base, and the one number that lands?
Locked
The failure mode of enterprise agents is organizational, not intellectual: they get cancelled for cost, value, and risk-control failures — exactly the layer CoWork's harness supplies.
>40%
of agentic AI projects will be canceled by the end of 2027 — not for lack of model intelligence, but for escalating costs, unclear business value, and inadequate risk controls. Source: Gartner press release, 25 June 2025. Caveat: a forecast, not an observed rate.

Supporting evidence

  • fact Worldwide AI spending forecast at $2.59T in 2026, +47% YoY (Gartner, May 2026) — the market is moving, and fast.
  • fact ~39% of IT time on custom integrations (MuleSoft, 2025) — the coordination problem never went away.

Contradictory / honest

  • The 40% figure is a prediction, not measured; treat it as directional, and say so on stage.
  • Incumbents argue software companies will absorb AI rather than be eaten — a real counter-thesis we must answer, not ignore.
  • open A precise dollar cost for "coordination" remains unquantified; we proxy it, we do not fake it.
4

Hydrate the Status Quo

What happens today, and why does it feel good enough?
Locked
The dominant workaround is manual coordination across point tools — and it feels "good enough" only because the pain lives at the seams between tools, not inside any one of them. Inertia is the real competitor.

Today's landscape

  • Do nothing: migration sits in the backlog for months.
  • Manual: platform teams hand-debug dozens of repos; coordinators shuttle across teams.
  • Point SaaS: CI/CD, identity, approvals, audit — each strong alone, none spanning the boundary.
  • Emerging AI: single-developer coding agents that are local, repo-scoped, and ungoverned.

Why anyone would change

  • A hard deadline (CVE, contract migration) is the moment "good enough" collapses.
  • The switch is from human coordination to agent execution with human steering.
  • Point tools still win inside a single function; they lose at the boundary between functions.
5

Hydrate the Opportunity & Insight

What changed, and why is the insight non-obvious?
Locked
The insight is not "AI + dev tooling." It is that the organizational layer — roles, consent, audit — can itself be re-embodied as agent infrastructure. SaaS's real product was that layer, and that layer is now buildable as a harness.

Problem → insight → opportunity

  • Problem: cross-boundary work is throttled by coordination/governance.
  • Insight: the thing SaaS sold was organizational problem-solving; now it maps onto a runtime.
  • Opportunity: the next enterprise product is a harness that makes organizational work agent-executable.

Why now

  • Agent capability crossed a threshold at the same moment enterprises hit a governance wall (the 40% cancellation).
  • The gap between "agents can" and "orgs will let them" is the market.
  • This ties directly to the theme: Strategy (what gets done), Processes (how it flows), Roles (humans steer, agents execute).
6

Hydrate the Competitive Landscape

Who else is here, and why haven't they already filled the gap?
Locked
The gap CoWork owns is the organizational harness as a first-class product — identity, consent, and audit built for agents, not bolted on to a human-facing UI.

Who is adjacent

  • Agent platforms (Devin, Codex cloud, Copilot Workspace) — strong on single-repo/cloud coding, weak on cross-repo governance.
  • Cloud dev environments (Codespaces, Gitpod, Daytona) — compute/workspace, not consent/audit.
  • Orchestration frameworks (LangGraph, CrewAI, AutoGen) — graph plumbing, no enterprise trust surface.
  • Incumbents (ServiceNow, Atlassian, GitHub, AWS) — org layer exists, but it is human-UI-shaped, not agent-runtime-shaped.

Why the gap remains

  • Incumbents are incented to keep humans inside their UIs — that is their revenue model.
  • Framework vendors lack the enterprise trust surface to carry security sign-off.
  • open Specific competitor claims should be re-verified on each vendor's own site before pitching — done as categories here, not as audited claims.
7

Hydrate the AI & Agent Opportunity

What becomes possible because of AI, and what roles do agents take?
Locked
AI's role is transformative, not cosmetic — but only inside a harness. A fleet of agents can traverse 4 repos or 60 services and apply one human's architectural decision across all of them; without consent and audit, that same capability is unshippable.

Agent roles

  • Worker agents apply changes across repos/services.
  • Steerer (human) injects architectural judgment at exceptions.
  • ConsentBroker plays the approval role at policy gates.
  • AuditLogger plays the compliance role (signed thought chains, diffs, receipts).
  • InteractionSurface plays the verification role (visual proof).

How the work actually changes

  • From humans doing to humans steering and approving.
  • Identity is task-scoped (IdentityBroker); actions are policy-gated (ConsentBroker).
  • This is not "add a chatbot to a tool" — it changes who does the work and who is accountable.
8

Hydrate the Product

What is the simplest coherent experience, and the one thing it must do well?
Locked
Dispatch a cross-boundary change, watch it execute in an ephemeral sandbox, steer the exceptions, approve, and get a signed record of everything. The one thing it must do well is make that work visible, steerable, and auditable end-to-end — not just generate code faster.

Core journey

  • Dispatch → mount workspace → agent acts → human steers exceptions → propagate pattern → consent/approve → signed receipt → linked PRs.
  • Inputs: task spec / breaking-change description.
  • Outputs: verified diffs, a PR suite, an audit receipt.

User stories

  • As a platform lead, I want to dispatch a 60-service migration, so that a CVE is remediated in days, not months.
  • As a security owner, I want a signed, tamper-evident record of every agent action, so I can approve without reading every line.
  • As a tech lead, I want to steer one bespoke legacy pattern, so my judgment propagates to the rest of the fleet.
9

Hydrate the MVP & Demonstration

What is real, what is simulated, and what is the 30-second "aha"?
Locked
The MVP is a deterministic, offline-safe interactive demo: real UI and state transitions, simulated cloud substrate and live inference. The "magic moment" is one human steering decision turning a whole fleet green.

Real vs simulated

  • Real: 4-step stepper, 60-tile fleet matrix, steering drawer, consent modal, terminal stream, preview canvas.
  • Simulated: EFS/ECS mounts, real Docker, Slack webhook, LLM inference (deterministic engine).

Magic moment & failure modes

  • <30s "aha": 15 amber tiles flip green after one human steering decision propagates, capped by a signed consent card.
  • Breaks first with 10 real users: trust in accuracy (simulator isn't inference), then cost/latency, then data access.
  • Live-demo fallback: the engine is fully deterministic and offline — the demo cannot go down.
10

Hydrate Validation

What assumptions could kill this, and what is the fastest credible test?
Locked
The kill-the-idea assumption is enterprise trust: if platform/security leads still refuse agents even with a signed audit trail, the harness thesis fails — no amount of demo polish saves it.

Riskiest assumptions

  • 1. Enterprises will trust agents with cross-boundary actions given consent + audit (vs. refusing outright).
  • 2. "Steer once, propagate N" is a real, generalizable pattern, not a demo artifact.
  • 3. The harness is defensible against incumbents adding governance later.

Fastest credible tests

  • 5–8 platform/security-lead interviews with a walking skeleton: would you approve a real change with this receipt?
  • Run propagation on 3–5 real open-source repos with one bespoke pattern each.
  • Competitor teardown: verify incumbent governance claims on their own sites.
11

Hydrate Business & Scale Potential

Who pays, for what, and is the beachhead deliberate?
Locked
Who pays is platform/security leadership, for risk reduction and delivery speed. The beachhead is deliberate: enterprise platform/security teams (they feel the CVE deadline first), expanding to any cross-system organizational work.

Economic logic

  • Value measured: time-to-remediation, fewer broken-staging incidents, faster audit pass-through.
  • Defensibility: the consent + audit + identity substrate is hard to bolt on; incumbents are incented against it.
  • Universal vs local: cross-boundary work is universal; the beachhead is a stated narrowing with a path outward.

Not forced

  • hypothesis Likely pricing is usage (sandbox-minutes / agent-hours) plus a governance seat — but this is a hypothesis, not a decision.
  • We do not force a business model just to look complete; the product thesis must land first.
12

Build Readiness Check

Is this hydrated enough to enter a two-week build?
Locked
Yes — with two named gaps. Problem, audience, insight, AI role, and MVP are hydrated enough to build. Enterprise trust and live competitor claims remain to be validated, not invented.

Ready to build

  • Problem, audience, trigger, insight, AI role, product, and MVP are each reducible to one sentence.
  • The 30-second demo is concrete and offline-safe.
  • The two scenarios are anchored as proof of the thesis, not as the pitch itself.

Gaps before full build

  • Validate enterprise trust with 5–8 real interviews.
  • Re-verify competitor claims on vendor sites.
  • Prove "steer once, propagate N" generalizes beyond the demo.
The proof, not the pitch

Two scenarios, one thesis

The existing scenarios stop being two separate demos and become two instances of the same claim: AI is eating the organizational problem-solving layer, and CoWork is the harness that lets it.

Scenario 1 — Contract migration

A breaking event schema change across 4 repos.

Today it needs 4 teams, shared staging that breaks for 50 engineers, and days of triage. CoWork mounts all four repos in one ephemeral sandbox, the agent refactors across them, boots the stack, and verifies the checkout flow end-to-end.

What's being eaten: the cross-team coordination and staging ritual — the organizational glue, not the code editor.

Beat: "No broken staging. Zero local setup. Four linked, pre-verified PRs."

Scenario 2 — Fleet migration

A CVE forces an SDK upgrade across 60 downstream services.

Today it takes months because each repo has bespoke legacy patterns. CoWork dispatches 60 workers; 45 pass, 15 pause. One human injects a compatibility shim on worker #18, propagates the pattern to the remaining 14, and all pass.

What's being eaten: the platform team's manual debugging across backlogs — the organizational bottleneck, not the SDK.

Beat: "Steer once, propagate to the fleet, sign off with a cryptographic receipt."
Final gate · Part 3 + 4

Build Readiness Record

Hydrated — ready for a two-week build

The idea is now a coherent, evidence-backed, buildable proposition rather than a longer one-pager. The checklist below stays answerable in one sentence per line:

  • The Problem — cross-boundary engineering work is throttled by coordination and governance, not by code-writing.
  • The Audience — platform/security leads, triggered by a breaking change or CVE that can't wait on backlogs.
  • The Core Value — turn organizational problem-solving into agent-executable infrastructure: steer once, scale everywhere, with a signed record.
  • The Mechanics — agents act in a scoped sandbox; humans steer and approve; brokers enforce identity, consent, and audit.
  • The Constraints — two weeks, offline-safe deterministic demo, beachhead on enterprise platform/security with a stated expansion path.

Most likely to kill the idea: enterprise teams still refuse autonomous cross-boundary agents even with a signed audit trail. That is the assumption to test first, before polishing the demo.

Sources

  1. Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027," 25 June 2025 — gartner.com. Forecast, not observed rate.
  2. Gartner, "Gartner Forecasts Worldwide AI Spending to Grow 47% in 2026," 19 May 2026 — gartner.com.
  3. MuleSoft, "10th Annual Connectivity Benchmark Report," 2025 — ~39% of IT team time on building/testing custom integrations (vendor survey of IT leaders).
  4. Internal: CoWork scenarios & MVP design spec and the interactive showcase.