AI governancea systematic literature review

Amna Batool, Didar Zowghi, Muneera Bano•View original
OverviewBalancedmaya & adam hosts
Maya: A company deploys an artificial intelligence hiring tool. A hospital starts using an artificial intelligence diagnostic system. A city rolls out predictive policing. Different sectors have different stakes, but the same question hangs over all three: when something goes wrong, who exactly is responsible? Adam: That question turns out to be much harder to answer than you might think. A team of researchers, including Amna Batool, Didar Zowghi, and Muneera Bano, spent considerable effort trying to map how much of the literature actually addresses it. Their systematic literature review analyzed twenty-eight studies on artificial intelligence governance to ask four questions: who is accountable, what elements need governing, when governance happens in the development life cycle, and how it is implemented. Maya: The paper is valuable. I want to say that upfront. But I also want to emphasize something it reveals about itself: mapping the territory is not the same as charting a route through it. The map they produce is honest, and part of its honesty is showing where the roads are not built yet. Adam: Start with why this matters at all. The authors note that recorded artificial intelligence incident repositories already contain over three thousand events. Autonomous vehicles, healthcare systems, and financial algorithms — these aren't hypothetical risks. The European Union Artificial Intelligence Act, the Organisation for Economic Co-operation and Development, the National Institute of Standards and Technology, the International Organization for Standardization, and the Institute of Electrical and Electronics Engineers — major institutions have all issued principles and guidelines. So the problem isn't a shortage of concern. Maya: The problem is fragmentation. Responsibility is dispersed across developers, deployers, regulators, and users, often with no clear line between them. The paper distinguishes ethical artificial intelligence, which adheres to moral principles around fairness, privacy, and human dignity, from responsible artificial intelligence, which focuses on development and deployment practices that minimize bias and discrimination. These are complementary aims, the authors say, both oriented toward building trust. However, knowing what you're aiming at and knowing who is supposed to aim are different things. Adam: Walk me through how they constructed the review. Maya: They searched Scopus and Google Scholar using the terms "AI," "artificial intelligence," and "governance," which were deliberately narrow to reduce noise. The initial pool contained nearly two thousand nine hundred records. Abstract screening reduced that to two hundred and twenty-five. Full-text review narrowed it down to sixty-one candidates. Then, two systematic literature review experts performed three rounds of selection, using forward and backward snowballing. This process involves checking both what a paper cites and what cites it, which expanded and then refined the set to the final twenty-eight. Adam: Quality was assessed using a five-question checklist from Liu and colleagues, scoring each paper on a scale from zero to one. Twenty of the twenty-eight were rated as Good. Eight were rated as Fair. None were rated as Poor. Maya: Which is a reasonable corpus. The limitation I would point out — and the authors acknowledge it themselves — is scope. The search covered journal databases only. There was no grey literature, no ISO standards, and no documents from the National Institute of Standards and Technology. The authors plan to extend the review to include those sources, explicitly naming ISO and NIST AI risk frameworks as targets. Adam: That limitation matters more than it may seem, because a lot of governance actually exists in those standards documents. Maya: Now to the findings — and this is where it gets interesting. The who and when questions are the least answered in the literature. Of twenty-eight studies, only six explicitly state who is accountable for artificial intelligence governance. Only three studies — Zhang and Zhang, Liao and colleagues, and Sun and Medaglia — answered all four questions: who, what, when, and how. That's three out of twenty-eight. Adam: That is a striking gap. Maya: The stakeholder categories that do appear are varied. These include artificial intelligence ethics committees convened by governments, a proposed Global Governance Coordinating Committee for international coordination, multi-disciplinary steering committees at the organizational level, hospital administrators, public managers, and state members of the World Health Organization and the International Telecommunication Union as international overseers. The roles exist, but the assignment of responsibility between them largely does not. Adam: On timing, or the when, eight studies argue that governance needs to address specific stages of the artificial intelligence life cycle. The authors organize these stages as pre-development, during development, and post-development. Five of those eight call for governance that spans all three stages. One study highlighted the importance of human involvement during development for population health applications. Another specifically recommended focusing governance efforts at the pre-development and post-deployment stages to reduce clinical ethical risks. Maya: The pattern shows that what to govern and how to govern it generate far more literature than who bears the burden and when they bear it. Nine studies proposed concrete governance artifacts, including frameworks, models, tools, and policies. The ethical principles most frequently targeted were fairness first and privacy second. There is more material on the substance of governance than on its enforcement or its timing. Adam: The five-level taxonomy is where the what and how findings get organized: team, organization, industry, national, and international. Seven of the nine concrete governance solutions sit at the organizational level. Maya: Seven out of nine. Adam: The organizational artifacts are the most concrete. ECCOLA, or Ethical Considerations and Challenges of Learning Algorithms, is a tool-based intervention at the organizational level, according to study A25. Study A19 extends it to align with Generally Accepted Recordkeeping principles. These are things a team can actually pick up and use. Maya: At the industry level, the Dimensional AI Governance model from study A10 covers structural, relational, and procedural components for market contexts. A Media-AI Governance Framework from study A1 targets risks in media specifically and is organized around four automation layers: data capture and processing, content generation, content moderation, and communication. Adam: International level? Maya: One governance solution from study A12 spans design-phase policies, testing and validation policies, deployment policies, and monitoring and oversight procedures. This includes the proposed Global Governance Coordinating Committee from study A3 and the World Health Organization and International Telecommunication Union Focus Group for Healthcare AI from study A23. These are mechanisms, but they are thin on implementation. Adam: Team level and national level governance solutions? None. There are zero concrete implemented artifacts at either level in the reviewed set. Maya: That's the gap that matters most for practitioners. If you're a developer on a team trying to figure out how to govern your specific AI project day to day, this literature largely doesn't serve you. At the organizational level, there are tools you can use. However, at the team level, there is almost nothing. Adam: The regional variation finding presents a different perspective. European work emphasizes human rights and data protection, particularly through the General Data Protection Regulation and the European Union Artificial Intelligence Act. In contrast, the United States favors market-driven, sectoral approaches and voluntary guidance provided by the National Institute of Standards and Technology. The Asia-Pacific region is diverse: China adopts a government-driven approach, while Singapore and Japan emphasize responsible, human-centric methods. Australia takes a stance that combines ethics with innovation. These differences are not merely stylistic; they reveal fundamentally different assumptions about where accountability should reside. Maya: This highlights the international coordination gap even further. Only one study offers an international-level governance solution, which is a policies and procedures framework rather than a binding mechanism. Adam: Here's what I keep coming back to. The authors write that their findings "can assist research communities in proposing comprehensive AI governance practices." Research communities, not practitioners. The paper is explicit about this, and I think that's the honest read. Maya: It is the honest read. The taxonomy across five levels, the World Health Organization, what, when, and how framework, and the gap analysis — these are inputs for building the next generation of governance tools. Three out of twenty-eight studies answered all four questions. The largest concentration of solutions sits at one level. Ethical guidelines frequently remain suggestions without enforceable regulations. Those findings tell researchers where to work. Adam: The diagnostic is real. The gap between a diagnosis and a treatment plan is also real, and the paper does not close it — and to be fair, it does not claim to. Maya: The next step mentioned by the authors is the extension into grey literature and standards, including ISO and NIST documents. This includes practice-level material that didn't make it into the first pass. That work is apparently ongoing. Adam: As AI deployment accelerates, the governance infrastructure is catching up, but it is uneven. The organizational level is the most developed, while the team level and international coordination are the least. The questions of who and when are lagging behind the what and how by a significant margin in the literature. If you are trying to decide which governance framework fits your AI system, this paper outlines what the options look like and highlights where the shelf is nearly empty. That is genuinely useful. It's just not the same as stocking the shelf. Maya: This means the work ahead focuses precisely on those empty spots — team-level tools that practitioners can actually use, international mechanisms with some enforcement capability, and governance artifacts that specify not just what to do but also who is responsible and when in the development process they are supposed to do it. Adam: The map exists. The roads still need to be built. Maya: This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
A

A company deploys an artificial intelligence hiring tool. A hospital starts using an artificial intelligence diagnostic system. A city rolls out predictive policing. Different sectors have different stakes, but the same question hangs over all three: when something goes wrong, who exactly is responsible?

B

That question turns out to be much harder to answer than you might think. A team of researchers, including Amna Batool, Didar Zowghi, and Muneera Bano, spent considerable effort trying to map how much of the literature actually addresses it. Their systematic literature review analyzed twenty-eight studies on artificial intelligence governance to ask four questions: who is accountable, what elements need governing, when governance happens in the development life cycle, and how it is implemented.

A

The paper is valuable. I want to say that upfront. But I also want to emphasize something it reveals about itself: mapping the territory is not the same as charting a route through it. The map they produce is honest, and part of its honesty is showing where the roads are not built yet.

B

Start with why this matters at all. The authors note that recorded artificial intelligence incident repositories already contain over three thousand events. Autonomous vehicles, healthcare systems, and financial algorithms — these aren't hypothetical risks. The European Union Artificial Intelligence Act, the Organisation for Economic Co-operation and Development, the National Institute of Standards and Technology, the International Organization for Standardization, and the Institute of Electrical and Electronics Engineers — major institutions have all issued principles and guidelines. So the problem isn't a shortage of concern.

A

The problem is fragmentation. Responsibility is dispersed across developers, deployers, regulators, and users, often with no clear line between them. The paper distinguishes ethical artificial intelligence, which adheres to moral principles around fairness, privacy, and human dignity, from responsible artificial intelligence, which focuses on development and deployment practices that minimize bias and discrimination. These are complementary aims, the authors say, both oriented toward building trust. However, knowing what you're aiming at and knowing who is supposed to aim are different things.

B

Walk me through how they constructed the review.

A

They searched Scopus and Google Scholar using the terms "AI," "artificial intelligence," and "governance," which were deliberately narrow to reduce noise. The initial pool contained nearly two thousand nine hundred records. Abstract screening reduced that to two hundred and twenty-five. Full-text review narrowed it down to sixty-one candidates. Then, two systematic literature review experts performed three rounds of selection, using forward and backward snowballing. This process involves checking both what a paper cites and what cites it, which expanded and then refined the set to the final twenty-eight.

B

Quality was assessed using a five-question checklist from Liu and colleagues, scoring each paper on a scale from zero to one. Twenty of the twenty-eight were rated as Good. Eight were rated as Fair. None were rated as Poor.

A

Which is a reasonable corpus. The limitation I would point out — and the authors acknowledge it themselves — is scope. The search covered journal databases only. There was no grey literature, no ISO standards, and no documents from the National Institute of Standards and Technology. The authors plan to extend the review to include those sources, explicitly naming ISO and NIST AI risk frameworks as targets.

B

That limitation matters more than it may seem, because a lot of governance actually exists in those standards documents.

A

Now to the findings — and this is where it gets interesting. The who and when questions are the least answered in the literature. Of twenty-eight studies, only six explicitly state who is accountable for artificial intelligence governance. Only three studies — Zhang and Zhang, Liao and colleagues, and Sun and Medaglia — answered all four questions: who, what, when, and how. That's three out of twenty-eight.

B

That is a striking gap.

A

The stakeholder categories that do appear are varied. These include artificial intelligence ethics committees convened by governments, a proposed Global Governance Coordinating Committee for international coordination, multi-disciplinary steering committees at the organizational level, hospital administrators, public managers, and state members of the World Health Organization and the International Telecommunication Union as international overseers. The roles exist, but the assignment of responsibility between them largely does not.

B

On timing, or the when, eight studies argue that governance needs to address specific stages of the artificial intelligence life cycle. The authors organize these stages as pre-development, during development, and post-development. Five of those eight call for governance that spans all three stages. One study highlighted the importance of human involvement during development for population health applications. Another specifically recommended focusing governance efforts at the pre-development and post-deployment stages to reduce clinical ethical risks.

A

The pattern shows that what to govern and how to govern it generate far more literature than who bears the burden and when they bear it. Nine studies proposed concrete governance artifacts, including frameworks, models, tools, and policies. The ethical principles most frequently targeted were fairness first and privacy second. There is more material on the substance of governance than on its enforcement or its timing.

B

The five-level taxonomy is where the what and how findings get organized: team, organization, industry, national, and international. Seven of the nine concrete governance solutions sit at the organizational level.

A

Seven out of nine.

B

The organizational artifacts are the most concrete. ECCOLA, or Ethical Considerations and Challenges of Learning Algorithms, is a tool-based intervention at the organizational level, according to study A25. Study A19 extends it to align with Generally Accepted Recordkeeping principles. These are things a team can actually pick up and use.

A

At the industry level, the Dimensional AI Governance model from study A10 covers structural, relational, and procedural components for market contexts. A Media-AI Governance Framework from study A1 targets risks in media specifically and is organized around four automation layers: data capture and processing, content generation, content moderation, and communication.

B

International level?

A

One governance solution from study A12 spans design-phase policies, testing and validation policies, deployment policies, and monitoring and oversight procedures. This includes the proposed Global Governance Coordinating Committee from study A3 and the World Health Organization and International Telecommunication Union Focus Group for Healthcare AI from study A23. These are mechanisms, but they are thin on implementation.

B

Team level and national level governance solutions? None. There are zero concrete implemented artifacts at either level in the reviewed set.

A

That's the gap that matters most for practitioners. If you're a developer on a team trying to figure out how to govern your specific AI project day to day, this literature largely doesn't serve you. At the organizational level, there are tools you can use. However, at the team level, there is almost nothing.

B

The regional variation finding presents a different perspective. European work emphasizes human rights and data protection, particularly through the General Data Protection Regulation and the European Union Artificial Intelligence Act. In contrast, the United States favors market-driven, sectoral approaches and voluntary guidance provided by the National Institute of Standards and Technology. The Asia-Pacific region is diverse: China adopts a government-driven approach, while Singapore and Japan emphasize responsible, human-centric methods. Australia takes a stance that combines ethics with innovation. These differences are not merely stylistic; they reveal fundamentally different assumptions about where accountability should reside.

A

This highlights the international coordination gap even further. Only one study offers an international-level governance solution, which is a policies and procedures framework rather than a binding mechanism.

B

Here's what I keep coming back to. The authors write that their findings "can assist research communities in proposing comprehensive AI governance practices." Research communities, not practitioners. The paper is explicit about this, and I think that's the honest read.

A

It is the honest read. The taxonomy across five levels, the World Health Organization, what, when, and how framework, and the gap analysis — these are inputs for building the next generation of governance tools. Three out of twenty-eight studies answered all four questions. The largest concentration of solutions sits at one level. Ethical guidelines frequently remain suggestions without enforceable regulations. Those findings tell researchers where to work.

B

The diagnostic is real. The gap between a diagnosis and a treatment plan is also real, and the paper does not close it — and to be fair, it does not claim to.

A

The next step mentioned by the authors is the extension into grey literature and standards, including ISO and NIST documents. This includes practice-level material that didn't make it into the first pass. That work is apparently ongoing.

B

As AI deployment accelerates, the governance infrastructure is catching up, but it is uneven. The organizational level is the most developed, while the team level and international coordination are the least. The questions of who and when are lagging behind the what and how by a significant margin in the literature. If you are trying to decide which governance framework fits your AI system, this paper outlines what the options look like and highlights where the shelf is nearly empty. That is genuinely useful. It's just not the same as stocking the shelf.

A

This means the work ahead focuses precisely on those empty spots — team-level tools that practitioners can actually use, international mechanisms with some enforcement capability, and governance artifacts that specify not just what to do but also who is responsible and when in the development process they are supposed to do it.

B

The map exists. The roads still need to be built.

A

This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Credits

  1. Source: AI governance: a systematic literature reviewAmna Batool, Didar Zowghi, Muneera BanoAI and Ethics, 2025DOI: 10.1007/s43681-024-00653-wCC-BY 4.0

AI-generated lecture by ennepō.ai, adapted from the source above — restructured, summarized and explained. Not endorsed by the authors.

This lecture © ennepō.ai. See our Terms of Service.

Report a problem with this lecture

More in Social Sciences →