Local Governments Need Evidence Before Scaling AI Tools

California is rolling out AI resources for local governments, but experts warn that workshops alone are insufficient. Agencies must prove that new workflows actually improve service delivery and safety before expanding usage.
A recent workshop at Pepperdine University highlighted a critical gap in how local governments approach artificial intelligence. While public-sector leaders are eager for hands-on practice with real-world scenarios, many are stopping at the classroom stage. This leaves staff with new skills but no clear path to integrate them into daily operations safely.
The urgency is high as California makes discounted AI tools and training available to cities and counties. However, access is expanding faster than most agencies can redesign their core processes. Without a structured approach, this rush risks creating confusion, inconsistent data handling, and tools that are either underused or employed without proper oversight.
Training Must Precede Operational Rules
Experts advise treating training as the starting point for a workflow test rather than the end goal. Before staff open any tool, managers must define specific operating rules. This includes identifying which systems are approved, listing data that cannot be entered, and assigning named reviewers for outputs. Vague instructions like "use your judgment" are insufficient; clear accountability structures are necessary to prevent unauthorized or unsafe usage.
The goal is to avoid a common pattern where employees attend a workshop, learn a few prompts, and return to offices where no approved tasks exist. Some staff stop using the tools entirely, while others use them quietly without guidance. A few may produce impressive demonstrations that never translate into reliable public services. Establishing rules first ensures that adoption is grounded in policy rather than experimentation.
Measuring Value Beyond Activity Metrics
To determine if an AI workflow is successful, agencies must compare the new process against existing methods. Metrics such as login counts or prompt usage indicate activity, not value. Instead, leaders should measure time saved, error rates, rework, and the overall experience for both employees and residents. A pilot that generates more drafts but requires more correction work has not improved government efficiency.
According to GN technics/ai (en-US), a responsible evaluation requires a brief after-action review. Workers should be encouraged to identify where the system helped, where it failed, and what they were reluctant to report. This feedback loop is crucial because employees may hide mistakes if leadership focuses too heavily on adoption rates. Creating psychological safety turns frontline staff into an effective early-warning system for potential issues.
Avoiding Vendor-Driven Adoption Pitfalls
This evidence-based approach protects smaller jurisdictions from being pushed into broad platform purchases by vendors. A county does not need to buy a comprehensive solution just because a demonstration looked impressive. By starting with a narrow problem, such as summarizing public comments or drafting routine communications, agencies can test the process and invest only when the evidence supports expansion.
Managers play a decisive role in this process by making time for practice and protecting employees who flag errors. AI use should not be treated as a performance target. If a responsible pilot concludes that human work remains faster or safer, that finding is still valuable. It prevents wasteful spending and ensures that technology serves the public interest rather than just checking a box.






