Spatial Is Special, Governing AI in the Geospatial Age



SHARE

As AI systems become ubiquitous, organizations are rushing to feed them every kind of data imaginable, including spatial data. According to Amy Rose, who works on data interoperability at Overture Maps, treating geospatial data like any other dataset in an AI pipeline is a recipe for unreliable results, unintended privacy violations, and decisions no one can audit.

We sat down with Rose to discuss what practical governance for spatial AI actually requires, who should be making the key decisions, and why the gap between “AI-ready” and “AI-trustworthy” remains stubbornly wide.

Why Spatial Really Is Special

For years, Rose was among those in the geospatial community who argued that “spatial is not special” and that isolating spatial data from mainstream technology only limits its reach. However, when it comes to AI governance, she has come to a different conclusion: spatial really is special.

“When you’re thinking about non-spatial governance, you’re dealing with rows and columns,” Rose explains. “But when you’re dealing with spatial data, now you have things that are inherently part of the spatial domain, like scale, precision, proximity, boundaries, and derived location. All of those things add layers of complexity that most people using AI simply aren’t accounting for.”

The core problem, she says, is that users often assume the model has knowledge of the data. When location data is introduced into an AI system, users make assumptions about fitness for use, data provenance, and the validity of combining different datasets—assumptions that may or may not be correct. Worse, there is often no auditability: you get an answer, but there is no way to trace it back to the source.

Building a Governance Framework That Works

A practical governance framework for spatial AI, Rose argues, must go beyond access control. It needs to account for the entire decision chain, not just who can see the data, but how data from different sources is combined, what purpose each dataset was intended for, and what limitations apply.

“It has to include where the data came from, what the intended purpose is supposed to be, what the limitations are, anything from license to access restrictions,” she says. 

“You have many situations where people have access to data for different purposes, and AI may or may not be aware of what that means.”

Rose offers a concrete example: an organization might have legitimate access to high-resolution imagery, insurance policy data, and property records each for a separate, approved purpose. If someone then asks an AI system to cross-reference those datasets to identify properties that should have their insurance policies changed, they are combining data in ways that could expose personally identifiable information (PII) and using it for purposes it was never intended for. On top of the privacy risk, they may also be introducing errors into their results.

“It’s not just about who has access,” Rose says. “It has to involve a combined decision-making process.” It’s not governance of a model; it’s governance of that decision chain and understanding the provenance, the lineage, and every aspect of the data that came into it, as well as being able to audit the output.”

Who Decides What’s In Scope?

When asked who should be deciding which spatial tables and attributes are in scope for a given AI-driven question, Rose is clear: these decisions should never be left to implementation time. By the time a model is being deployed or a question is being asked, the governance framework should already be in place.

The decision should involve data owners, she says, and should be structured around a decentralized model rather than a single owner controlling everything. Metadata, access controls, and provenance information should all be attached to the data itself.

The harder question is never the technical one. “We all can agree, ‘Hey, this would be a good technical solution,'” Rose says. “It’s about when it comes time to actually make those decisions.” Those decisions often need to be made quickly, especially in crises like natural disasters or conflicts, where the pressure to bring data together fast is intense. In those situations, an imperfect answer is often considered better than no answer at all.

The danger, Rose warns, is in the chain of custody. Someone may produce an answer with full knowledge of its caveats and provenance, but as that answer gets passed along—analyzed, turned into a graph, dropped into a PowerPoint—the context gets stripped away. “At some point, you’re really losing the essence of how you got to that answer.”

In practice, Rose says, the state of governance is “all over the board.” In higher-risk domains like insurance, some organizations are beginning to put structure around the decision chain to ensure the right people are involved and the process is auditable. By and large, there is a lot of open space with no actual governance mechanism in place.

This gap is part of why Overture Maps is explicitly building lineage and provenance into the data it produces, Rose notes. “If nothing else, that is part of the data that we are producing.”

The Biggest Gap: Trusting the Output

The biggest gap between organizations that look AI-ready and those that can actually trust their AI spatial outputs comes down to a fundamental misunderstanding of how AI works with spatial data.

LLMs are trained on text, Rose explains. They recognize language patterns. They do not have native geospatial capabilities. “You could ask, ‘What country is Paris located in?’ and it’ll probably tell you France, but not because it’s doing a spatial intersection. It’s because it’s seen ‘Paris, France’ a trillion times in text and recognizes that pattern.”

This means that people who aren’t familiar with geospatial technology may be asking spatial questions without understanding how those questions are being answered. The system may be doing pattern matching rather than actual location computation. AI, as Rose notes, is “obviously famous for hallucinating”; it will give you an answer regardless of whether that answer is correct.

The solution, she argues, is straightforward but often overlooked: humans absolutely must remain in the loop. “AI can’t be a black box. Do you understand the process? If you no longer had AI and had to do this manually, do you understand the steps that it took? And are those the right steps?”

The ubiquity and ease of tools like ChatGPT and Claude is a double-edged sword. Making AI that accessible is great, but in the geospatial world, and with data in general, the easiest path is not always the one that produces a trustworthy, accurate answer.

Entity Resolution: The Hidden Bottleneck

Before closing, Rose raised a point she felt hadn’t been asked about: the challenge of entity resolution. Beyond the data itself, the interoperability problem, how different datasets work together, is the bigger challenge Overture Maps is trying to solve.

Overture’s Global Entity Reference System is often described as a way to solve the “conflation” tax—making it easier to bring data together. But Rose sees its real significance in entity resolution, especially as AI and eventually quantum computing tackle increasingly complex optimization problems.

“If I’m a county GIS person versus somebody who holds the tax records versus somebody who’s working for a construction company, all these different people have a piece of information about the same building,” Rose says. “Right now, there is literally no way for them to easily understand whether they’re all talking about the same building.”

Without fast, reliable entity resolution, that ambiguity will remain an incredible limiting factor in the world of AI, or worse, it will produce a lot of bad data output. The future of trustworthy spatial AI, it turns out, depends not just on better models or better formats, but on the unglamorous, foundational work of knowing whether two datasets are actually talking about the same thing.

TAGS

About the Author

View more from

Related

Most Read