Germany faces a structural administrative problem—and the federal government believes it has found the solution: AI in public administration, not as a digital writing assistant, but as an autonomous colleague. Under the banner “Colleague AI,” the newly created Federal Ministry for Digital Affairs and State Modernization (BMDS)—led by Digital Minister Karsten Wildberger (CDU)—is working to introduce agent-based AI systems into government agencies that review applications, analyze documents, and provide decision recommendations—while humans formally retain the final say. The plan sounds ambitious, and the first pilot projects are already underway. But behind the facade of press releases and promises of acceleration lurk structural issues that have hardly been discussed publicly so far.
What the “Agentic AI Hub” really is—and what it promises
In February 2026, the Federal Ministry for Digital Affairs and State Modernization (BMDS) launched the so-called “Agentic AI Hub.” According to BMDS, 18 pilot projects were selected from around 400 startups and 200 interested municipalities (an independent verification of these figures is not available at the time of writing), which are now being tested in real government agencies. The Federal Digital Service is supporting the implementation of these projects. According to a ministry press release, State Secretary Thomas Jarzombek (BMDS) articulated the ambition as follows: “We want to build a ramp for startups into the public administration.”
The selected projects demonstrate just how concrete the vision already is—though the following examples are based on the companies’ own statements and are still in the early pilot phase. Forml is working in Frankfurt and Düsseldorf on the automated verification of housing eligibility certificates—a process that has previously been manual and time-consuming. Formfix supports the processing of long-term care applications in Cologne. In the Neckar-Odenwald district, Lector.ai processes incoming government mail using a Vision-LLM that recognizes and categorizes scanned documents. These projects are not just a digital gimmick—they intervene in real administrative processes that directly affect people’s lives.
Wildberger himself promotes the approach with big numbers: AI agents could “speed up approval procedures by over 80 percent.” The figure sounds impressive. What’s missing is the methodology behind it—no benchmark, no reference process, no independent basis for evaluation. This is no trivial matter, but a symptom: the political pressure to show results is greater than the willingness to measure them accurately.
The demographic pressure behind it: Why government agencies are looking into this at all
To understand why the federal government is taking this course, one must know the context. Germany is facing a massive staffing shortage in the public sector. The baby boomer generation will be retiring in the coming years, and there are not enough young people to fill the resulting gaps. At the same time, administrative burdens are growing—more funding programs, more complex EU regulations, and higher expectations from citizens regarding digital services.
By international standards, Germany is already under pressure. In the EU index for the digitization of public services, Germany consistently ranks near the bottom. Estonia is often cited as the counterexample: with the X-Road system, this small Baltic country has built a functioning digital administrative infrastructure that can now support AI services without having to start from scratch. Austria has at least established a digital records management system with ELAK (electronic file), on which AI pilots in federal ministries are already being built. Switzerland is proceeding more cautiously: The Federal Chancellery is testing AI tools with an explicit focus on GDPR and DSG compliance before anything is introduced into production systems.
It should be noted, however, that Estonia, with 1.3 million inhabitants and a historically developed digital infrastructure (e-Estonia since the 1990s), is hardly directly comparable to Germany. Administrative federalism, the tradition of data protection, and sheer size make Germany a structurally different case.
Germany, on the other hand, is embarking on the AI leap from a foundation that is not even fully digitized in many areas. Some municipalities still send notices by mail because the fax machine is the most modern means of communication available. Deploying AI agents designed to act autonomously on this ground is a technological gamble with a significant risk of failure.
AI in public administration and the EU AI Act: What “high risk” really means
That the topic of AI in public administration is not purely technical, but deeply legal, becomes clear at the latest in the context of the EU AI Act. The European regulatory framework, which will gradually come into force starting in 2024, classifies AI systems used for application review in government agencies as high-risk applications—and for good reason.
High-risk AI under the EU AI Act is subject to strict requirements: The systems must be explainable, meaning they must be able to provide a comprehensible rationale for why they arrive at a particular assessment. They must maintain complete logs. They must structurally enable human oversight—not just formally. And they must not produce discriminatory outputs, meaning they must not systematically disadvantage any population groups.
Now the honest question: Do the current pilot projects already meet these requirements? The answer is unclear — and that is the problem. Neither the BMDS nor the participating startups have yet disclosed in detail how their systems handle explainability and bias control. Green Party MP Rebecca Lenhard has publicly criticized precisely this: There is a lack of transparency and evaluation standards. Who checks whether a pilot project is actually better than the previous process after six months—and according to what criteria?
The answer to this question is not merely academic. It will determine whether the “Agentic AI Hub” goes down in history in two years as a success story or as a costly experiment.
The structural limitations: bias, black box, and automation bias
Behind the political debates lie three technical and social problems that receive too little attention in the public discussion.
The black box problem: Many of the most powerful AI systems—particularly large language models, on which solutions like Lector.ai or similar systems are based—are not fully transparent in their decision-making logic. It is often impossible to reconstruct step by step why one application is classified as complete and another is not. For a government agency that, according to the Basic Law, must act neutrally, objectively, and in accordance with the law, this is a central challenge. A decision based on an incomprehensible AI recommendation is legally vulnerable—according to prevailing opinion, an unsubstantiated AI-supported decision is likely to violate the requirement to state reasons (§ 39 VwVfG) and be subject to challenge in an appeal proceeding.
The bias problem: AI systems learn from historical data. If this data reflects past inequalities—such as certain population groups having received negative decisions more frequently in the past—then an AI system can perpetuate and reinforce these patterns. A housing eligibility certificate system that systematically rates applications from certain neighborhoods more poorly would be a nightmare for administrative law and social cohesion.
Automation Bias: Even if a human formally makes the final decision, research shows that people tend to uncritically adopt machine recommendations—especially under time pressure and when dealing with a high caseload. “The computer suggested it” becomes the de facto decision, even if the case worker formally signs off on it. This effect is well documented in aviation, medicine, and the judiciary. It will be no different in public administration.
The principle that “humans make the decision” is therefore necessary but not sufficient. Active structures are needed—training, dual-control principles, regular audits—to ensure that human oversight does not degenerate into a mere formality.
The Other Perspective: What Tech Optimists Object
It would be incomplete to focus solely on the risks. A growing group of administrative reformers and technologists argues that the risks are manageable—and that the costs of inaction are higher than the risks of action. In this scenario, explainability and transparency are overestimated: People make unexplainable decisions in government agencies every day without this being considered a “fundamental problem.”
Added to this is the pragmatic position: Full automation is not the goal. Even 30–50% automation of routine tasks noticeably reduces the workload for case workers—without every decision being made autonomously. The risk profile of semi-automated processing differs from that of a fully autonomous AI decision. Pilot projects such as the housing eligibility certificate review show: When AI pre-sorts and humans decide, the majority of legal issues are eliminated. Perfect systems are not a prerequisite for useful systems.
What is needed now: standards, transparency, and the courage to take things slowly
The 18 pilot projects of the Agentic AI Hub are a first step. But for pilots to truly become scalable systems, more is needed than press releases and promises of acceleration.
First, binding evaluation standards are needed. Who evaluates the pilot projects? According to which metrics—processing time, error rate, equal treatment of different groups, citizen satisfaction? These standards must be defined before the pilot, not after. Independent oversight bodies—such as the Federal Commissioner for Data Protection and Freedom of Information in collaboration with technical experts—should have access to the system data.
Second, a realistic infrastructure is needed. AI tools for public administration can only be as good as the data they operate on. As long as records exist in different systems—some still on paper—the degree of automation is inevitably limited, and the risk of errors is high. Investing in basic data infrastructure is less glamorous than an AI pilot project, but it is the prerequisite for everything else.
Third, we need the courage to proceed slowly in the right places. The vision of “nationwide scalability” sounds good in a press release. But 18 pilot projects in heterogeneous municipal environments are not a sufficient foundation for a national rollout—not for high-risk applications that determine housing, care services, and other essential goods. It took Estonia 30 years to build its digital administration. Germany will not catch up in two years.
That doesn’t mean AI agents in public administration are wrong. The staff shortage is real, the pressure to act is real, and the technology is indeed powerful enough to deliver real benefits. But the difference between “powerful” and “legally compliant” is not a technical footnote in public administration—it is the crux of the matter.
Well-known projects such as BärGPT in Berlin or LLMoin at Dataport in Hamburg show that there is a middle ground: internal language models for research and text drafting, with clear boundaries on what can and cannot be decided automatically. This is less spectacular than autonomous agents—but it is an approach that is legally sound and politically justifiable.
The real question is not whether AI should be used in German public administration. It has long been here. The question is how the transition will be managed—and whether political impatience will override the necessary methodological care before citizens bear the consequences.
🎼
The Doctor’s Opinion
AI agents in public administration are no longer a technical problem—they are a governance problem. The BMDS pilot projects are, in principle, the right approach, but 80-percent promises without methodology and evaluation standards without independence are warning signs, not proof of success. Anyone who truly wants “Colleague AI” must first take the infrastructure, legal compliance, and control functions seriously—otherwise, the hoped-for accelerator will turn into an expensive boomerang.