Government agencies must be held responsible for negative outcomes from their use of AI
ID 303143856 © Tero Vesalainen | Dreamstime.com

Commentary

Government agencies must be held responsible for negative outcomes from their use of AI

Government agencies must pair artificial intelligence deployment with clear rules, human oversight, and enforceable accountability for internal and vendor use.

Government agencies are increasingly using artificial intelligence (AI) to reduce administrative burdens and improve public services. When agencies apply AI to practical, well-defined tasks, it can speed up routine work. This can free public employees to focus on complex cases and provide direct services.

However, AI can also produce errors or unfair outcomes when government agencies use it to influence decisions about people’s rights, benefits, or safety without adequate testing and oversight. For example, a system used to screen benefit applications could incorrectly identify a person’s application as suspicious or incomplete, delaying assistance while the case is reviewed. A law enforcement system could incorrectly connect a person or vehicle to an investigation, leading officers to question or investigate someone who is not involved. These actions can also be difficult for residents and agency staff to understand or correct. AI can magnify these existing government failures by applying the same flawed rules, inaccurate data, or biased patterns across many cases in a short period. 

AI can also create privacy and security risks when government agencies or vendors collect, retain, share, or expose sensitive government information without adequate controls. A breach or unauthorized use of agency data can reveal personal information or allow data to be used beyond legal or agency-approved limits.

To minimize these harms, government agencies must pair AI deployment with clear rules, human oversight, and enforceable accountability for agencies and vendors. 

Streamlining government work

AI can help public employees complete repetitive, high-volume tasks more quickly. For example, it can prepare a first draft of a routine letter, summarize an authorized case file, or pull information from a standardized form. Agencies should begin with defined, low-risk tasks and measure whether a tool is accurate, secure, and worth its cost before expanding its use. 

Governments are already showing how AI can support routine work at every level of the public sector. The Food and Drug Administration (FDA) uses its internal Elsa platform to assist with scientific review and operational tasks. Elsa helps staff review clinical protocols and summarize reports of harmful side effects. It can also search large document collections, compare labels,  analyze data and generate code. Subject-matter experts verify the inputs, analytic processes, and implementation of AI-supported outputs. 

The North Carolina Department of State Treasurer uses ChatGPT in its Unclaimed Property Division and related offices. The tool helped employees locate businesses that may hold unclaimed assets, summarize regulations and lengthy materials, analyze financial information, and draft clearer communications for businesses and residents. Pilot participants reported saving 30 to 60 or more minutes each day. 

Cities are also moving beyond small-scale testing. San Francisco expanded Microsoft 365 Copilot Chat from a pilot of more than 2,000 employees to nearly 30,000 staff members. The platform helps employees analyze data, assist with project planning, and conduct research. For public servants such as social workers, these time savings can mean more time for direct services. By sharing results and lessons from these pilots, governments can build on effective practices and avoid repeating mistakes as they expand AI use.

Testing and correcting AI errors

Before a government agency uses AI to support decisions about benefits, rights, employment, education, healthcare, or access to government services, it should test whether the system is accurate for the task. The agency should compare its results with reliable information and check for patterns of error across communities and case types. For example, before using AI to flag benefit applications or generate law enforcement leads, agencies should compare its results with actual outcomes and check for false flags.

Agencies should not treat AI recommendations as a final decision or treat an AI-generated match, risk score, or investigative lead as proof that a person committed wrongdoing or is connected to an investigation. A trained employee should be required to review the relevant facts, be able to disagree with the system, and document the reason for a significant decision.  

Agencies should also monitor systems after deployment, investigate recurring errors, and pause or change a tool if it produces unreliable or unfair results. They should track how often staff overrides the system, whether errors cluster in particular types of cases or among particular communities, and whether residents are experiencing delays, denials, or other harmful outcomes. 

Protecting sensitive data and government records

Government agencies hold sensitive information about their residents, including health, tax, criminal justice, personnel, and investigative records. AI does not change the government’s duty to protect that information. It requires agencies to apply existing privacy, security, confidentiality, and records-management obligations to new tools, vendors, and workflows. These obligations apply to other government and contractors as well, but AI deserves added attention because it can process and reuse large volumes of government information quickly, allowing a single weak control to expose more data or affect more cases. Data safeguards help prevent these harms by protecting sensitive information and preserving records that agencies may need to investigate misuse, security incidents, or problems with an AI system.

Before a government employee uses an AI tool, the agency should know where the information will go, who can access it, whether the vendor will keep it, and whether the vendor may use it for another purpose. Without those controls, an employee could enter sensitive case information into a public AI tool and expose it to a system the agency cannot audit, secure, or delete. Generative AI systems can also create privacy risks through the leakage, unauthorized disclosure, or de-anonymization of sensitive personal information.  

AI can create records that agencies must retain, manage, or disclose under existing public-records law. A prompt submitted to an AI tool, the tool’s chat histories, and its output may all qualify as government records depending on how they are used. For example, an AI-generated summary used to support an agency recommendation may document official business and need to be preserved. Agencies should therefore use only approved tools and prohibit staff from providing sensitive information to unapproved systems. Contracts should limit data retention, restrict access, and require deletion when the contracts end. Agencies should also use encryption to protect sensitive information, maintain audit logs to identify improper access or misuse, and establish procedures for responding to a breach or other failure. Federal guidance also cautions agencies to protect government information from unauthorized disclosure and to restrict vendors from using it to train or improve commercial products without agency permission. 

Accountability for government AI use

Government agencies must remain accountable for how they use AI and for the effects of AI-supported decisions.  Using AI does not excuse agencies from civil rights, anti-discrimination, privacy, or due-process obligations. When an AI-supported action affects benefits, employment, licensing, healthcare, education, or public safety, residents should be able to seek an explanation, correction, and meaningful human review. 

AI can also change the scale of agency action. A mistake in a vendor’s system, an inaccurate AI model, or an access setting that lets the wrong people see or use the data can affect many records, decisions, or residents at once. Agencies need controls that can detect and stop these failures before they spread.

Vendor-operated systems can also create accountability problems. For example, automated license-plate-reader systems allow police departments and other public agencies to collect and search vehicle and location information at scale. Agencies set the rules for access, retention, sharing, and auditing, but a vendor’s system design and configuration can determine whether those rules are carried out as intended. 

Concerns about broad data sharing, potential access by federal immigration authorities, and improper searches have led some jurisdictions to reconsider or cancel contracts for automated license-plate-reader systems. In Ventura, Calif., the police department found that a vendor-based configuration error had allowed unauthorized out-of-state law enforcement agencies to query local license-plate-reader data, even though the department’s settings were intended to limit access to California partners. The city said the queries occurred without its knowledge or authorization. 

The Ventura incident shows why government contracts cannot simply assume that a vendor’s technical settings reflect an agency’s legal and policy limits. Agencies should set rules for data access and sharing, then use audit logs to regularly test whether those rules are being followed. Contracts should require vendors to configure systems to enforce approved access limits, provide usable audit logs, promptly disclose misconfigurations and breaches, and cooperate with independent audits. They should also define who owns the data, prohibit unauthorized commercial reuse, and allow agencies to suspend or terminate a contract when safeguards fail. This would give agencies a clear way to stop using a system that does not meet required protections.

Government must always be accountable for how it uses its power, including when it uses AI. AI can improve services and reduce administrative burdens at scale, but it can also magnify errors, privacy violations, and surveillance. Strong safeguards and enforceable oversight can help ensure that governmental use of AI serves the public without allowing the government to evade responsibility for the harms it causes.