Technology & AIAnalysis

What Would Success Look Like for Japan's Government AI Experiment?

Japan's government has launched a large-scale test of organisational AI adoption — and said in advance what it will measure. Adoption and transformation are not the same question.

Eveil’s earlier analysis of AI adoption in Japan argued that the harder problem is not access to the technology but what happens once a working experiment forces an organisation to decide who owns the changed workflow, what data the system may see, and who is accountable for scaling it. That argument was built from private-sector survey data, where the honest answer to most of those questions was: nobody had asked yet.

Japan’s government is now running a version of that experiment at an unusually large organisational scale. It has also said, in public, before the results were in, what it intends to measure — and, in the same breath, what that measurement will not yet be able to tell anyone.

A test at government scale

From May 2026, the Digital Agency opened a large-scale demonstration of its government AI system, Gennai, to roughly 180,000 officials across every ministry and agency2. By the end of that month, about 100,000 already had access, with the remainder following through the year2. By August, the platform offered more than forty applications built for named administrative tasks — drafting Diet responses, searching legal and regulatory material, answering questions against internal systems — running against government-procured datasets that include nearly eighty years of the Official Gazette3.

The stated purpose is not simply to put a tool in front of officials. The Digital Agency’s own framing is quality improvement and efficiency in the work itself, opening room for more creative and strategic work rather than routine drafting and searching3. Part of the same effort has been opened outward: sections of Gennai were released as open-source software, with a further release under discussion as of August 2026, extending the platform toward local governments and industry4.

Scale is what makes the claim worth watching rather than merely stating. A pilot inside one ministry proves little about an organisation this size. A demonstration spanning every ministry is a genuine test of whether adoption at scale produces anything an outside observer would recognise as transformation.

Adoption is not the same question as transformation

Speaking on 17 April, several weeks before the demonstration began, Japan’s Digital Minister set out how the government meant to judge it. Distributing accounts, he said, is not the goal in itself — what matters is verifying whether Gennai is actually being used, and whether that use is actually useful1. Toward that verification, the Digital Agency planned to publish monthly figures from around summer 2026: a utilisation-start rate, the cumulative share of account-holders who have opened Gennai at least once; a monthly utilisation rate, the share using it at least once within a given month; and executions per user1.

Those are sensible things to measure, and it is worth saying so plainly. An organisation with no adoption cannot have transformation — a licence nobody opens changes nothing about how work gets done. Publishing the figures by ministry also creates a comparison that did not previously exist, which is itself a form of pressure toward use.

But the three metrics answer a narrower question than the headline suggests. A utilisation-start rate establishes that someone opened the tool once. A monthly rate establishes that someone returns to it. Neither says what they used it for, whether the output was any good, or whether anything downstream of the click actually changed.

Eveil View Adoption metrics can tell us whether people are using AI. They cannot tell us whether the organisation is transforming. Utilisation rates and executions per user can show that use is spreading and deepening — a materially easier question than whether the work got better, or whether anything about how it is done has actually changed.

What makes the April statement unusual is what came next. Asked directly how effect evaluation was progressing, the Minister did not reach for a number. Measuring the effect itself, he said, is the actual purpose of the demonstration going forward — meaning it has not happened yet. The most accessible proxy available is likely to be staff surveys, because a specific figure for hours saved on a Diet response would need pre-adoption data for comparison, and that baseline probably does not exist1.

That is a government stating, ahead of its own results, that usage numbers and outcome evidence are different things, and that only the first is ready.

The Minister’s answer also points at a distinct problem, one worth separating from the metrics themselves. What to track once a tool is rolled out — the utilisation and execution figures above — is one question. What baseline needs to exist before rollout, if an organisation later wants to show that the work itself improved, is a different one. The two are easy to conflate, because both look like measurement, but only the second can ever answer whether performance changed.

That exposes a measurement problem larger than Gennai. If the baseline is not established before scaling, an organisation may later be able to show that AI was widely used without being able to show how much the work improved because of it — not because the improvement did not happen, but because nothing was measured that could prove it either way. The gap the Minister described is not a shortcoming particular to this demonstration; it is what tends to happen when adoption is scaled faster than the measurement designed to evaluate it.

This is not the first time Gennai’s use has been measured this way. A year before the government-wide demonstration began, the Digital Agency’s own roughly 1,200 staff were already being tracked on the same kind of metric: about 80% had used the tool at least once over three months, and total use exceeded 65,000 instances5. The aggregate number concealed a wide spread. Some 184 staff used it more than a hundred times each, while 178 used it fewer than five times, and roughly half of section-chief-level staff recorded no use at all5. An adoption figure that looks healthy in aggregate can still describe a workforce where use is concentrated in a minority — a caution worth carrying into any reading of the 180,000-official figures still to come.

What harder evidence would look like

If usage figures are a floor rather than an answer, it is worth being precise about what a stronger case would need to show. Access asks whether employees can reach the capability at all — largely settled here. Usage asks whether they return to it for work that matters, rather than trying it once and stopping. Productivity asks whether it measurably reduces time, cost or duplicated effort. Quality asks whether the output, the analysis or the decision it supports is actually better. And changes in work asks the hardest thing of all: whether workflows, role boundaries, handoffs or who is accountable for what have genuinely moved.

This is deliberately not a maturity model, and treating it as one would overstate what the categories can do. Each is a different kind of evidence, not a rung on a ladder every organisation climbs in the same order. What the ordering does show is where the Digital Agency’s committed metrics actually sit: almost entirely inside the first two categories. That is not a criticism of the choice — usage is where any credible measurement effort has to start. It is a statement of how much is still open.

Why the coming agent environment raises the stakes

The government has already said what comes next for the platform, even before outcome evidence exists for its current use. Around October 2026, Gennai is due to add an environment in which officials build their own AI agents rather than only using applications the Digital Agency supplies3. An employee describes a task in natural language; the system turns that description into a “skill” — a structured document capturing the procedure, the judgement rules, the reference material and the required output format — which then runs as an agent, chaining tens or hundreds of steps to carry out the task3. A marketplace is planned alongside it, so skills can be shared, evaluated and reused across ministries: one official’s accumulated know-how, converted into what the Digital Agency describes as the government’s collective knowledge3.

That single design choice turns Part I’s operating-model questions from abstract into immediate. Once employees can create agents, “who owns the changed workflow” becomes “who is allowed to build one, and for what.” “What data may the system access” becomes a question asked at the moment of creation, not discovered after deployment. Quality assurance, version control, and who answers for it when a shared skill produces a wrong result inside another ministry’s process are not implementation footnotes.

They are the operating model, arriving earlier and faster than most organisations plan for.

What remains uncertain The material reviewed here does not describe how skill quality will be assured before an agent reaches the marketplace, what happens when a reused skill is wrong, or how accountability follows a skill once another ministry’s employee has downloaded and run it. Those are precisely the questions Part I identified as the difficult part of an AI transformation. Government scale gives an unusual opportunity to observe how they are addressed in practice — not a guarantee that they will be addressed, or addressed visibly.

The broader Japan question

Gennai sits inside a wider pattern. IPA’s most recent national survey — 1,799 companies, fielded mid-April to mid-June 2026 — found just under a third of firms already using AI reporting effects at or above expectation, and roughly half reporting some effect that fell short6. Where the effect showed up mattered more than whether it did: 91.6% reported work becoming more efficient or faster, against 4.5% reporting better customer satisfaction and 3.9% reporting higher revenue6. The report’s own subtitle asks the question directly — AI adoption is spreading; can DX actually change6.

That is the same gap Gennai’s own metrics are exposed to. Efficiency and speed surface first because they sit closest to the tool; whether decisions, workflows or accountability have actually moved is slower to show and harder to see — the difference between adoption and transformation, in a ministry or a private company.

What the experiment can and cannot yet tell us

Gennai is worth watching not primarily for its size, but because Japan’s government has said in advance what it will measure and, almost in the same sentence, said what that measurement will not yet show. That degree of candour about what the current metrics can and cannot show is itself worth noting.

The harder test has not been run. It cannot be, until usage has had time to become routine and the agent environment has had time to be used for something real. Whether Gennai succeeds or fails is not yet a question the evidence can answer, and treating this year’s usage figures as if it already had would be a misreading of what the government itself has said those figures mean.

What to validate next Before scaling any AI rollout — inside government or outside it — decide two things, not one: how adoption will be tracked, and what baseline needs to exist now if the organisation later wants to show the work actually changed. The first is usually easy to arrange after the fact. The second is not — a baseline missed before rollout may be difficult, and sometimes impossible, to reconstruct afterward. An organisation that can only name the first has confirmed that AI is being used. It has not yet given itself a way to know whether that use was worth anything.

  1. 松本大臣記者会見(令和8年4月17日)デジタル庁Government / regulatorJapanese source
  2. (参考資料)ガバメントAI源内の展開状況デジタル庁 戦略・組織グループ AI実装総括班Government / regulatorJapanese source
  3. 政府内における生成AI「源内」の活用について(AI時代の国家公務員の在り方に関する検討会 第1回 資料5)デジタル庁 源内班Government / regulatorJapanese source
  4. 「ガバメントAI源内」OSS Ver 2.0の計画に関するオンライン説明会の開催についてDigital AgencyGovernment / regulator
  5. デジタル庁職員による生成AIの利用実績デジタル庁 戦略・組織グループ AI実装戦略総括Government / regulatorJapanese source
  6. DX動向2026 — 広がるAI導入、DXは変われるか独立行政法人情報処理推進機構 (IPA)Government / regulatorJapanese source
Read more on this theme
Technology & AI
Need this answered for your own situation
Intelligence — structured work on one management question, rather than public analysis of a market.
Need to decide what to do about it
Advisory — working through the decision itself.