AI in Public GovernanceChallenges and Future Directions

OverviewBalancedjames voice
Picture a government office somewhere in Europe. An algorithm is running quietly in the background, flagging citizens it suspects of benefit fraud. It processes thousands of cases a day, far more than any human team could manage. Then it gets one wrong. Not just wrong, but wrong in a way that ruins someone's year, maybe longer. When the question comes — who is responsible, how was this decision made, can you explain it — the answer is silence. That gap between the algorithm's confidence and the institution's accountability is exactly what two major research reviews have spent years trying to understand. The promise is real, and both teams take it seriously. Babšek and colleagues, working from a Scopus dataset of three thousand one hundred and forty-nine documents, frame the opportunity for public administration across three dimensions: internal processes, service delivery, and policymaking. On the internal side, that means document processing, automated decision support, predictive maintenance of infrastructure, and big-data analysis. For citizens, it means chatbots, personalized services, public-health monitoring, and emergency response. For policymakers, artificial intelligence offers predictive forecasting, risk assessment, and evidence-based modeling of policy outcomes. Zuiderwijk, Chen, and Salem synthesize a complementary catalog from their review of twenty-six studies — nine categories of benefit, from efficiency and economic gains to better decision-making and sustainability. Taken together, the picture is genuinely attractive: faster internal operations, more responsive services, and richer policy insights. Both teams also note we are in a period of high expectations around artificial intelligence. This is exactly why the obstacles matter as much as they do. Before getting to those obstacles, it's worth understanding how small and selective the evidence base actually is. Babšek et al. started with three thousand one hundred and forty-nine documents and concentrated on the top two hundred most-cited articles by citations per year — roughly five percent of the initial set, using a threshold of about ten citations per year. After excluding sixty-three general reviews, they conducted in-depth content analysis of one hundred and thirty-seven papers. Zuiderwijk et al. took a different path: a systematic review protocol that started with eighty-five unique studies, screened out forty-eight after reading titles and abstracts, and then excluded eleven more after full-text quality assessment. That left twenty-six studies. All twenty-six were published within the three years prior to their review, twenty-one used qualitative methods, and only two were quantitative. Both exercises reveal a field that is young, concentrated, and largely qualitative. The meta-finding is important: we don't yet have a broad, empirical base for many of the claims made about artificial intelligence in public administration. That caveat shadows everything that follows. The first and most foundational challenge is data — not the algorithms, not the hardware, but the raw material that feeds everything else. Zuiderwijk et al. identify data quality and data sharing as distinct, persistent obstacles. Government data tends to live in siloed agencies with legacy systems, inconsistent formats, and strict privacy constraints that limit what can be shared, combined, or fed into a model. Public-sector data is structurally different from the kind of clean, centralized data that private companies build artificial intelligence on. Babšek et al. reinforce this: limited or poor-quality government data appears repeatedly as a barrier across their top-cited literature. The practical consequence is that even a well-designed artificial intelligence system can fail before it runs a single prediction, because the inputs it needs simply aren't available in a usable form. Data problems aren't a technical inconvenience to be cleaned up later; they're a structural feature of how public institutions have historically managed information. Even when the data is good, the organization receiving the artificial intelligence often isn't ready. Both papers are clear that the biggest barriers to adoption are human and institutional. Babšek et al. identify conceptual ambiguity as a distinct and underappreciated problem: public administrators frequently lack a shared understanding of what "artificial intelligence" even means. Without that shared vocabulary, coherent strategy, procurement, and oversight become nearly impossible. You can't govern what you can't define. Zuiderwijk et al. add the skills dimension — government agencies lack the blended managerial and technical expertise needed to deploy artificial intelligence responsibly. Financial pressures compound the problem: building the necessary infrastructure and competing for a limited pool of artificial intelligence talent is expensive, and public institutions rarely win that competition against the private sector. Both reviews describe a chain of institutional failure: no shared concept of artificial intelligence, no appropriate skills, no governance routines, and no strategic alignment from leadership. Under those conditions, even capable models applied to good data will struggle to become reliable, legitimate public services. Then there is the governance gap — and this is where the two studies converge most forcefully. Zuiderwijk et al. frame the ethical and accountability challenges as core governance crises, not peripheral technicalities. They ask directly how artificial intelligence implementation affects accountability when government officials make decisions based on algorithmic outputs. They note that artificial intelligence must be transparent to a meaningful degree in order to retain citizens' confidence. The legal and policy frameworks governing public administration simply have not kept pace with artificial intelligence's development. Babšek et al. make that problem visceral with concrete cases. The Dutch SyRI system — an artificial intelligence tool for detecting child-benefit fraud — was criticized for privacy violations and lack of transparency. A Polish artificial intelligence system for classifying the unemployed drew similar fire. A UK visa-screening tool was abandoned over bias and discrimination concerns. Austria's AMS employment-profiling system faced legal scrutiny for reinforcing socioeconomic inequalities. In France, artificial intelligence-driven security initiatives, including a suspicious-noise detector and facial-recognition trials in high schools, were ruled unlawful by data-protection authorities. These are not edge cases; they are the documented track record of public-sector artificial intelligence deployment in Europe, drawn from the same literature Babšek et al. analyze. What unifies these failures is the interpretability problem. When an algorithm influences whether someone receives a benefit or gets flagged as a fraud risk, someone has to be able to explain why — and someone has to be responsible when it's wrong. Babšek et al. are direct: "the question of liability and responsibility in case of faulty or wrong decisions made by artificial intelligence technology is not clarified yet." Zuiderwijk et al. underscore that this ambiguity is especially dangerous in public governance, where decisions carry public-law consequences and where citizens have rights that algorithms don't automatically respect. Some jurisdictions are beginning to respond — mandatory algorithmic impact assessments, explainability requirements, and the European Union's artificial intelligence act with its risk-level categorizations — but Babšek et al. place these emerging remedies against a landscape where controversial deployments have already done reputational damage. Both papers close by pointing toward the same destination, though from different angles. Zuiderwijk et al. propose eight process-related and fifteen content-related recommendations in total, and the core message is that the field needs to mature methodologically. Only four of their twenty-six studies even cite a theory. They call for more empirical, explanatory research; quantitative and mixed-method approaches; better data sharing among researchers; and multidisciplinary foundations that draw on digital government, information systems, and related fields. They also flag a terminological problem that parallels Babšek et al.'s conceptual ambiguity finding: nine of the twenty-six studies treat artificial intelligence as a single undifferentiated phenomenon rather than studying a specific technology. Studying "artificial intelligence in government" without specifying whether you mean machine learning, natural language processing, or computer vision is like studying "medicine" without specifying the disease. Babšek et al. push the complementary, practice-facing side of that agenda. They triangulate their literature findings with real-world examples from the European Commission's Public Sector Tech Watch Observatory and treat the European Union's regulatory architecture as an emerging model for what responsible deployment looks like. Their emphasis is on concrete implementation metrics, rigorous evaluation of public value, and grounded use cases that can be transferred across jurisdictions. What neither paper fully resolves — and both implicitly acknowledge — is that many of the hardest obstacles are political and structural, not technical. Data silos persist because agencies are incentivized to protect their information. Skills gaps persist because government can't match private-sector compensation. Accountability gaps persist because defining legal liability for algorithmic decisions requires political choices that legislatures have been slow to make. Better algorithms and better research matter enormously. But Zuiderwijk et al. and Babšek et al. together make the case that closing the gap between artificial intelligence's promise and its reality in public administration will take something harder to engineer than a better model — it will take institutional will. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Picture a government office somewhere in Europe. An algorithm is running quietly in the background, flagging citizens it suspects of benefit fraud. It processes thousands of cases a day, far more than any human team could manage. Then it gets one wrong. Not just wrong, but wrong in a way that ruins someone's year, maybe longer. When the question comes — who is responsible, how was this decision made, can you explain it — the answer is silence. That gap between the algorithm's confidence and the institution's accountability is exactly what two major research reviews have spent years trying to understand. The promise is real, and both teams take it seriously. Babšek and colleagues, working from a Scopus dataset of three thousand one hundred and forty-nine documents, frame the opportunity for public administration across three dimensions: internal processes, service delivery, and policymaking. On the internal side, that means document processing, automated decision support, predictive maintenance of infrastructure, and big-data analysis.

For citizens, it means chatbots, personalized services, public-health monitoring, and emergency response. For policymakers, artificial intelligence offers predictive forecasting, risk assessment, and evidence-based modeling of policy outcomes. Zuiderwijk, Chen, and Salem synthesize a complementary catalog from their review of twenty-six studies — nine categories of benefit, from efficiency and economic gains to better decision-making and sustainability. Taken together, the picture is genuinely attractive: faster internal operations, more responsive services, and richer policy insights. Both teams also note we are in a period of high expectations around artificial intelligence. This is exactly why the obstacles matter as much as they do. Before getting to those obstacles, it's worth understanding how small and selective the evidence base actually is. Babšek et al. started with three thousand one hundred and forty-nine documents and concentrated on the top two hundred most-cited articles by citations per year — roughly five percent of the initial set, using a threshold of about ten citations per year. After excluding sixty-three general reviews, they conducted in-depth content analysis of one hundred and thirty-seven papers.

Zuiderwijk et al. took a different path: a systematic review protocol that started with eighty-five unique studies, screened out forty-eight after reading titles and abstracts, and then excluded eleven more after full-text quality assessment. That left twenty-six studies. All twenty-six were published within the three years prior to their review, twenty-one used qualitative methods, and only two were quantitative. Both exercises reveal a field that is young, concentrated, and largely qualitative. The meta-finding is important: we don't yet have a broad, empirical base for many of the claims made about artificial intelligence in public administration. That caveat shadows everything that follows. The first and most foundational challenge is data — not the algorithms, not the hardware, but the raw material that feeds everything else. Zuiderwijk et al. identify data quality and data sharing as distinct, persistent obstacles. Government data tends to live in siloed agencies with legacy systems, inconsistent formats, and strict privacy constraints that limit what can be shared, combined, or fed into a model.

Public-sector data is structurally different from the kind of clean, centralized data that private companies build artificial intelligence on. Babšek et al. reinforce this: limited or poor-quality government data appears repeatedly as a barrier across their top-cited literature. The practical consequence is that even a well-designed artificial intelligence system can fail before it runs a single prediction, because the inputs it needs simply aren't available in a usable form. Data problems aren't a technical inconvenience to be cleaned up later; they're a structural feature of how public institutions have historically managed information. Even when the data is good, the organization receiving the artificial intelligence often isn't ready. Both papers are clear that the biggest barriers to adoption are human and institutional. Babšek et al. identify conceptual ambiguity as a distinct and underappreciated problem: public administrators frequently lack a shared understanding of what "artificial intelligence" even means. Without that shared vocabulary, coherent strategy, procurement, and oversight become nearly impossible. You can't govern what you can't define. Zuiderwijk et al. add the skills dimension — government agencies lack the blended managerial and technical expertise needed to deploy artificial intelligence responsibly.

Financial pressures compound the problem: building the necessary infrastructure and competing for a limited pool of artificial intelligence talent is expensive, and public institutions rarely win that competition against the private sector. Both reviews describe a chain of institutional failure: no shared concept of artificial intelligence, no appropriate skills, no governance routines, and no strategic alignment from leadership. Under those conditions, even capable models applied to good data will struggle to become reliable, legitimate public services. Then there is the governance gap — and this is where the two studies converge most forcefully. Zuiderwijk et al. frame the ethical and accountability challenges as core governance crises, not peripheral technicalities. They ask directly how artificial intelligence implementation affects accountability when government officials make decisions based on algorithmic outputs. They note that artificial intelligence must be transparent to a meaningful degree in order to retain citizens' confidence. The legal and policy frameworks governing public administration simply have not kept pace with artificial intelligence's development. Babšek et al. make that problem visceral with concrete cases. The Dutch SyRI system — an artificial intelligence tool for detecting child-benefit fraud — was criticized for privacy violations and lack of transparency. A Polish artificial intelligence system for classifying the unemployed drew similar fire.

A UK visa-screening tool was abandoned over bias and discrimination concerns. Austria's AMS employment-profiling system faced legal scrutiny for reinforcing socioeconomic inequalities. In France, artificial intelligence-driven security initiatives, including a suspicious-noise detector and facial-recognition trials in high schools, were ruled unlawful by data-protection authorities. These are not edge cases; they are the documented track record of public-sector artificial intelligence deployment in Europe, drawn from the same literature Babšek et al. analyze. What unifies these failures is the interpretability problem. When an algorithm influences whether someone receives a benefit or gets flagged as a fraud risk, someone has to be able to explain why — and someone has to be responsible when it's wrong. Babšek et al. are direct: "the question of liability and responsibility in case of faulty or wrong decisions made by artificial intelligence technology is not clarified yet." Zuiderwijk et al. underscore that this ambiguity is especially dangerous in public governance, where decisions carry public-law consequences and where citizens have rights that algorithms don't automatically respect.

Some jurisdictions are beginning to respond — mandatory algorithmic impact assessments, explainability requirements, and the European Union's artificial intelligence act with its risk-level categorizations — but Babšek et al. place these emerging remedies against a landscape where controversial deployments have already done reputational damage. Both papers close by pointing toward the same destination, though from different angles. Zuiderwijk et al. propose eight process-related and fifteen content-related recommendations in total, and the core message is that the field needs to mature methodologically. Only four of their twenty-six studies even cite a theory. They call for more empirical, explanatory research; quantitative and mixed-method approaches; better data sharing among researchers; and multidisciplinary foundations that draw on digital government, information systems, and related fields. They also flag a terminological problem that parallels Babšek et al.'s conceptual ambiguity finding: nine of the twenty-six studies treat artificial intelligence as a single undifferentiated phenomenon rather than studying a specific technology. Studying "artificial intelligence in government" without specifying whether you mean machine learning, natural language processing, or computer vision is like studying "medicine" without specifying the disease.

Babšek et al. push the complementary, practice-facing side of that agenda. They triangulate their literature findings with real-world examples from the European Commission's Public Sector Tech Watch Observatory and treat the European Union's regulatory architecture as an emerging model for what responsible deployment looks like. Their emphasis is on concrete implementation metrics, rigorous evaluation of public value, and grounded use cases that can be transferred across jurisdictions. What neither paper fully resolves — and both implicitly acknowledge — is that many of the hardest obstacles are political and structural, not technical. Data silos persist because agencies are incentivized to protect their information. Skills gaps persist because government can't match private-sector compensation. Accountability gaps persist because defining legal liability for algorithmic decisions requires political choices that legislatures have been slow to make. Better algorithms and better research matter enormously. But Zuiderwijk et al. and Babšek et al. together make the case that closing the gap between artificial intelligence's promise and its reality in public administration will take something harder to engineer than a better model — it will take institutional will. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Social Sciences