AI LabAgent Infrastructure
Omniscope: an intelligence layer over the order platform.
An enterprise order platform knows everything and tells you almost none of it without someone clicking through screens. Omniscope makes it askable: 34 tools across nine domains, every one a read.
- Role
- Built it end to end
- Works with
- Any MCP client
- Runs on
- Cloudflare Workers
- Access
- Read-only by construction
> Why is order 4471 stuck?⏺ omniscope · investigate_order(orderId: "4471", env: "production") ⎿ Verdict Held. The payment authorisation expired before release. Timeline Mon 09:12 Order created Mon 09:13 Payment authorised Thu 09:13 Authorisation expired Thu 11:40 Release attempted, order held Evidence ✓ payment.status = AUTH_EXPIRED (get_order_payments) ✓ order.status = Held (get_order) ✓ failed.messages = 0 (get_failed_messages)⏺ Order 4471 is held because its payment authorisation expired before the order was released. Nothing failed between systems. Re-authorise the payment and the release can go through. 3 reads · 0 writes · productionWhat an MCP server is.
MCP, the Model Context Protocol, is an open standard for giving an AI assistant tools it can use. Omniscope is one of those servers.
Think of it as USB-C for AI. Any assistant that speaks the protocol can plug into any server that speaks it, and the server decides what the assistant can see and do.
Every assistant needs its own integration with every system, and every one is another way in to secure.
- The client
- The AI app you talk to: Claude Code, Cursor, an internal assistant. It reads your question and decides which tool to call.
- The protocol
- A shared format. The server lists its tools with a name, a description and the inputs each takes; the client calls one and gets JSON back.
- The server
- Omniscope. It holds the credentials, decides which reads exist at all, and answers. The assistant never touches the platform directly.
One question, as it crosses the wire
Why is order 4471 stuck?tools/list → 34 toolstools/call investigate_order { "orderId": "4471" }{ verdict, timeline, evidence: [ tool, field, value ] }Why I built it
Every order runs through sixty-odd connected systems: orders, payments, inventory, a message bus, a scheduler. Nothing documents how they fit together. So every everyday question meant clicking through screens built for people, one at a time. The APIs could answer faster, but they can also change data, so no AI assistant could be trusted with them.
Why is this order stuck?
Before
Five screens, in the right order
Effort
With Omniscopeinvestigate_orderIs anything breaking between systems?
Before
A manual count, by message type
Effort
With Omniscopeintegration_healthDoes production still match staging?
Before
Two tabs and a careful eye
Effort
With Omniscopecompare_environmentsWhich endpoint does this button call?
Before
Devtools, and an afternoon
Effort
With Omniscopeget_browser_activity
So I built a safe way to ask the platform directly. Every answer was already there.
Nine domains, thirty-four tools.
What the server can read, grouped by the kind of question it answers.
11 tools, Kind of knowledge: Live state
Orders
Stuck, held and failed orders across every market, each with a verdict and its evidence.
3 tools, Kind of knowledge: Live state
Integrations
Failed messages between systems, counted by type and by how badly.
1 tool, Kind of knowledge: Live state
Inventory
Where a unit is and every movement that got it there.
2 tools, Kind of knowledge: Live state
Scheduled work
Which jobs were due, which ran, and which did not.
3 tools, Kind of knowledge: Definition
API surface
What each operation does, and whether it may be called at all.
5 tools, Kind of knowledge: Definition
Configuration
The settings behind a behaviour, their history, and drift between environments.
2 tools, Kind of knowledge: Definition
Architecture
The dependency graph, and the contracts that span components.
4 tools, Kind of knowledge: Observation
Screen capture
What a real screen calls, recorded from a browser session.
3 tools, Kind of knowledge: Observation
Usage
How the server itself is used, and where it fails.
How a read gets through.
Four layers sit between the agent and the platform. Each can only narrow what is possible.
One request
MCP client
01Authenticate
OAuth, rate limit, audit row
No anonymous calls
02Allowlist
Documented reads only
No write can resolve
03Sanitise
PINs, cards, customer fields
No reveal path
04Evidence
Source, environment, true total
No partial answer reads as complete
Control Tower gateway
By the time a request reaches a source system, everything it could possibly do is a read.
My role
I built it as the analyst doing the investigations, which is why the tools match the questions people ask.
- 01
Mapped the domains
Took each domain from the work itself: which reads answer which question people actually ask.
- 02
Designed the safety model
Read-only and an explicit allowlist from the first line of code, not guards added later.
- 03
Wrote the Worker
TypeScript on Cloudflare Workers, including every gate described above.
- 04
Built the capture extension
A companion browser extension and the pipeline behind it, so a screen can be asked what it calls.
- 05
Named the failure patterns
Recurring failures became named verdicts, and the tools were reshaped around how people asked.
What changed
34
Questions the platform can now answer directly
Questions that each meant a different screen and someone who knew where to look are now one question.
- Before
- After
- Whoever did it last
- The same answer for anyone
- A conclusion to trust
- A conclusion with its evidence
- Drift found by accident
- Drift asked for on purpose
0
Write paths exposed
Daily
Use across the operations team
Built with
- TypeScript
- Model Context Protocol
- Cloudflare Workers
- Cloudflare KV
- Cloudflare D1
- OAuth
- REST APIs & JSON
- BigQuery
- Claude Code