Case study · Public interest / open government data and AI agents
Government open data that an AI agent can find, query and cite without a human in the loop
- Model Context Protocol
- WebMCP
- Python
- DuckDB
- Cloudflare Pages
- Cloudflare D1
- Cloudflare R2
✓ Verified Review“Most government sources publish their data as a CSV drop and leave it there. The Commonwealth's own data policy says that data should be machine-readable, available through an API and kept up to date automatically, so it can improve services and grow the economy. We built publicdata.au to do that work, so the engineers and AI agents building applications for citizens can use the data the way it was released to be used.”
Agent tools, served identically over MCP and WebMCP
Rows an agent can filter and aggregate across eight live datasets
Government catalogue records an agent can search and vote on
File formats generated from each version
The challenge
The Commonwealth's 2015 data policy commits agencies to machine-readable data with API access, kept up to date in an automated way. Most Australian government data is still published as a CSV or Excel drop on one of ten portals, with columns renamed between releases and files replaced in place. An engineer building an application for citizens, or an AI agent answering a question for one, has to find, parse and clean the file before any real work starts.
What we built
One pipeline that keeps every release and builds the files, the query API and the agent tools from the same model. Agents get a remote MCP server and in-browser WebMCP tools with identical definitions, and every answer carries its version, licence and attribution. A licence gate holds any dataset whose licence cannot be confirmed.
- A remote MCP server at /mcp with seven tools, listed in the official MCP Registry and installable in Claude, VS Code and Cursor
- The same tools registered in the browser through WebMCP, defaulting to the dataset on the page and describing its fields
- One definition of every tool and API parameter, generated into both surfaces and held to them by a gate and parity tests
- Discovery files for agents: llms.txt, an Agentic Resource Discovery manifest, an MCP server card and a Markdown twin of every page
- A licence gate that holds any dataset whose licence it cannot confirm and states the reason on the backlog page
- Immutable dated versions in ten formats, each with a manifest, the publisher's own file and a row-level diff
The stack
A stateless Streamable HTTP MCP server on Cloudflare Pages Functions, published to the MCP Registry on each release; WebMCP tools registered through navigator.modelContext; tool and API text written once in api.json and generated into both surfaces; a query API over Cloudflare D1; content-addressed raw storage in R2; and pure serialisers for ten formats.
The outcome
At its build of 28 September 2026 the site served eight datasets and 3,658,026 rows, including the national road deaths database back to 1989, New South Wales recorded crime by council area and month, and six Queensland road crash series. All of it is reachable through seven agent tools on two surfaces. The MCP server is active in the official MCP Registry as au.publicdata/mcp, and each release republishes the registry entry only when what a client sees has changed. Each queryable dataset is also an MCP resource listing its fields, their types, the publisher's descriptions and the values they hold, so an agent can read a dataset's shape before it writes a query.
Where it went next
The register holds five more datasets in its backlog, including the ACNC charity register and the ABN bulk extract. Agents can vote for the next dataset through the same tools people use, and the votes decide what follows. Every version is kept, and its raw bytes and manifests sit in both versioned object storage and git, so an answer an agent cited against a dated version can be rebuilt from either copy.
This case study describes a live product designed and built by National Digital. Figures are read from the live site's build of 28 September 2026. The site is independent: no government agency runs it, funds it or has endorsed it.
Key Takeaways
Engineering open data for AI agents
- Agents get first-class tools: a registered MCP server and WebMCP on every page.Critical
The remote MCP server needs no key or account and is listed in the official MCP Registry. In a browser that exposes WebMCP, the same tools register on each page and default to the dataset shown, so an agent can answer a question about the data it is reading.
- Each tool is defined once, so the two surfaces cannot drift apart.Critical
Every tool, parameter and limit lives in one api.json file that the build writes into the MCP server, the page scripts, OpenAPI and llms.txt. A gate fails the deploy if a copy drifts, and a test runs each tool through both surfaces and requires identical answers.
- Every answer is citable, because it names its version and carries the publisher's attribution.Critical
Tool answers include the version, the licence, the attribution string and the query URL, and the server instructions tell agents to show all three. A dated version always returns the same bytes, so a figure an agent cites can still be checked years later.
- Agents read a dataset's shape before they query it.Important
Each queryable dataset is an MCP resource listing its fields, their types, the publisher's descriptions, their ranges and the values of small fields. On a dataset page the WebMCP tools carry the same detail, so an agent sees the exact values a filter needs.
- The licence is read on every run and gates the publish.Important
CC BY and CC0 data publishes with the attribution in the publisher's own words. A licence the agency has not specified is held until a person records evidence, and a no-derivatives licence is never published. Old versions stay up under the licence they were published under.
Built by National Digital, publicdata.au serves Australian government open data to AI agents through a registered MCP server and WebMCP tools defined once and tested for parity. Every answer names its version and attribution, and every release is kept as an immutable dated version.
Questions decision-makers ask about this build
Why do governments release open data in the first place?
What does publicdata.au add to what agencies already publish?
How does an AI agent use publicdata.au?
How do you keep the MCP server, the browser tools and the API documentation consistent?
Can an agent cite what it finds?
Why republish data a government already publishes?
How does it handle licences?
Is it an official government site?
What's next
Want AI agents to use your data correctly?
On publicdata.au, agents get an MCP server and in-browser tools that answer from versioned data and cite their source. If your organisation publishes data or services that agents will act on, see how we approach platform engineering and data analysis and insight platforms, part of our wider custom software development work.