Case study · Public interest / open government data and AI agents

Government open data that an AI agent can find, query and cite without a human in the loop

  • Model Context Protocol
  • WebMCP
  • Python
  • DuckDB
  • Cloudflare Pages
  • Cloudflare D1
  • Cloudflare R2
✓ Verified Review

“Most government sources publish their data as a CSV drop and leave it there. The Commonwealth's own data policy says that data should be machine-readable, available through an API and kept up to date automatically, so it can improve services and grow the economy. We built publicdata.au to do that work, so the engineers and AI agents building applications for citizens can use the data the way it was released to be used.”

Cameron Young
CEO at National Digital
7

Agent tools, served identically over MCP and WebMCP

3,658,026

Rows an agent can filter and aggregate across eight live datasets

135,793

Government catalogue records an agent can search and vote on

10

File formats generated from each version

The challenge

The Commonwealth's 2015 data policy commits agencies to machine-readable data with API access, kept up to date in an automated way. Most Australian government data is still published as a CSV or Excel drop on one of ten portals, with columns renamed between releases and files replaced in place. An engineer building an application for citizens, or an AI agent answering a question for one, has to find, parse and clean the file before any real work starts.

What we built

One pipeline that keeps every release and builds the files, the query API and the agent tools from the same model. Agents get a remote MCP server and in-browser WebMCP tools with identical definitions, and every answer carries its version, licence and attribution. A licence gate holds any dataset whose licence cannot be confirmed.

  • A remote MCP server at /mcp with seven tools, listed in the official MCP Registry and installable in Claude, VS Code and Cursor
  • The same tools registered in the browser through WebMCP, defaulting to the dataset on the page and describing its fields
  • One definition of every tool and API parameter, generated into both surfaces and held to them by a gate and parity tests
  • Discovery files for agents: llms.txt, an Agentic Resource Discovery manifest, an MCP server card and a Markdown twin of every page
  • A licence gate that holds any dataset whose licence it cannot confirm and states the reason on the backlog page
  • Immutable dated versions in ten formats, each with a manifest, the publisher's own file and a row-level diff

The stack

A stateless Streamable HTTP MCP server on Cloudflare Pages Functions, published to the MCP Registry on each release; WebMCP tools registered through navigator.modelContext; tool and API text written once in api.json and generated into both surfaces; a query API over Cloudflare D1; content-addressed raw storage in R2; and pure serialisers for ten formats.

AgentsMCP server, WebMCP tools
Query APID1, OpenAPI, rate-limited
BuildOne model, ten formats, one tool spec
Raw storePublisher bytes by hash in R2
Gates & fetchLicence, provenance, tool parity

The outcome

At its build of 28 September 2026 the site served eight datasets and 3,658,026 rows, including the national road deaths database back to 1989, New South Wales recorded crime by council area and month, and six Queensland road crash series. All of it is reachable through seven agent tools on two surfaces. The MCP server is active in the official MCP Registry as au.publicdata/mcp, and each release republishes the registry entry only when what a client sees has changed. Each queryable dataset is also an MCP resource listing its fields, their types, the publisher's descriptions and the values they hold, so an agent can read a dataset's shape before it writes a query.

Where it went next

The register holds five more datasets in its backlog, including the ACNC charity register and the ABN bulk extract. Agents can vote for the next dataset through the same tools people use, and the votes decide what follows. Every version is kept, and its raw bytes and manifests sit in both versioned object storage and git, so an answer an agent cited against a dated version can be rebuilt from either copy.

This case study describes a live product designed and built by National Digital. Figures are read from the live site's build of 28 September 2026. The site is independent: no government agency runs it, funds it or has endorsed it.

Key Takeaways

Engineering open data for AI agents

  • Agents get first-class tools: a registered MCP server and WebMCP on every page.Critical

    The remote MCP server needs no key or account and is listed in the official MCP Registry. In a browser that exposes WebMCP, the same tools register on each page and default to the dataset shown, so an agent can answer a question about the data it is reading.

  • Each tool is defined once, so the two surfaces cannot drift apart.Critical

    Every tool, parameter and limit lives in one api.json file that the build writes into the MCP server, the page scripts, OpenAPI and llms.txt. A gate fails the deploy if a copy drifts, and a test runs each tool through both surfaces and requires identical answers.

  • Every answer is citable, because it names its version and carries the publisher's attribution.Critical

    Tool answers include the version, the licence, the attribution string and the query URL, and the server instructions tell agents to show all three. A dated version always returns the same bytes, so a figure an agent cites can still be checked years later.

  • Agents read a dataset's shape before they query it.Important

    Each queryable dataset is an MCP resource listing its fields, their types, the publisher's descriptions, their ranges and the values of small fields. On a dataset page the WebMCP tools carry the same detail, so an agent sees the exact values a filter needs.

  • The licence is read on every run and gates the publish.Important

    CC BY and CC0 data publishes with the attribution in the publisher's own words. A licence the agency has not specified is held until a person records evidence, and a no-derivatives licence is never published. Old versions stay up under the licence they were published under.

Built by National Digital, publicdata.au serves Australian government open data to AI agents through a registered MCP server and WebMCP tools defined once and tested for parity. Every answer names its version and attribution, and every release is kept as an immutable dated version.

Questions decision-makers ask about this build

Why do governments release open data in the first place?
Governments release data so that it gets used. The Australian Government Public Data Policy Statement of 7 December 2015 calls government data a strategic national resource that holds considerable value for growing the economy, improving service delivery and transforming policy outcomes. It commits agencies to publish non-sensitive data in a machine-readable format with API access, kept up to date in an automated way. NSW's open data policy aims at better public services.
What does publicdata.au add to what agencies already publish?
It does the part of the policy that a CSV drop leaves undone. Each release becomes a typed, versioned dataset with a Table Schema, ten file formats, a query API with OpenAPI and an MCP server an AI agent can call. Engineers can build an application on it without writing a parser, and every answer carries the publisher's attribution and the version it came from, so the application can show where each figure came from.
How does an AI agent use publicdata.au?
It connects to the remote MCP server at publicdata.au/mcp, which needs no key and is listed in the official MCP Registry, or it uses the WebMCP tools a page registers in the browser. Seven tools cover searching datasets, reading fields, querying and counting rows across the whole table, comparing versions and voting for new data. Claude, VS Code and Cursor can add the server in one step.
How do you keep the MCP server, the browser tools and the API documentation consistent?
Every tool, parameter and limit is written once in api.json, and the build generates the MCP tool list, the page scripts, OpenAPI, llms.txt and the agents page from it. A gate fails the deploy if a generated copy drifts or a tool has no server implementation, and a test runs each tool through the MCP server and the browser and requires identical answers.
Can an agent cite what it finds?
Yes. Every tool answer names the version it came from and carries the licence, the publisher's attribution string and the query URL, and the server tells agents to show them. Every data file carries a cite string in its header. A dated version URL always returns the same bytes, so a figure an agent quotes today can be checked against the same data years from now.
Why republish data a government already publishes?
Portals publish the current file and replace it when the next release arrives, so a figure cited from last year's file can no longer be checked against its source. The site keeps each release as a dated version, adds formats such as Parquet and SQLite that the publisher does not offer, and serves all of it to agents. Every page carries the attribution the publisher's licence asks for.
How does it handle licences?
The licence is read on every run and decides whether the dataset can be published. CC BY and CC0 data is published with the publisher's own attribution wording. A licence left as 'specified by agency' is held until a person records evidence of what it permits, and a no-derivatives licence is never published. Two datasets are blocked today for that reason, and the backlog page says so.
Is it an official government site?
It is independent. National Digital runs it, and no government agency runs it, funds it or has endorsed it. Every page and every data payload says so, and a build gate fails the deploy if a payload is missing that statement, its attribution or its source hash.

What's next

Want AI agents to use your data correctly?

On publicdata.au, agents get an MCP server and in-browser tools that answer from versioned data and cite their source. If your organisation publishes data or services that agents will act on, see how we approach platform engineering and data analysis and insight platforms, part of our wider custom software development work.

Want a platform like this for your business?

Tell us what you're building. We'll come back within one business day with what we'd do first and what it would take.