I first came across the field of Geographic Information Systems (GIS) while studying for a Master's in Environmental Remote Sensing at Aberdeen University. We needed GIS to help visualize remotely sensed data and derived analytics. I quickly realized that one of the core activities of a GIS practitioner was converting data for use in different software platforms. This was either a fitness-for-purpose activity, a projection activity (does anyone remember the infamous “datum conflict” error? One of my top five, for sure), or perhaps a data quality step. The idea of data preparation, management, and transformation has always been central to the role of a GIS person.
As I grew in my career through the 2000s, this only became more obvious. The GIS role was a transactive data role, even if we were using regular expressions for some tabular pivot in Awk (because at that point ArcInfo could not do those more complex attribute manipulations) or building programmatic indices from Landsat images, we were taking raw geographic products and manipulating the data (pixels, points or polygons) into something else for consumption in another interface. That interface could have been a website, a GIS, a map, a report, a spreadsheet, or a PowerPoint. Very rarely was the data the final product. The data was the resource, and we would refine that resource into the product our customers or stakeholders were seeking.
Every step in that chain of events would be held in the analyst's head. And each step is akin to a “break in gauge.” This is a railway term indicating a change in the gauge of railway tracks, necessitating one of a variety of complex engineering operations, or passengers and freight changing trains entirely. For us, this is a point at which a data product transforms to meet another product standard.

The idea of a “break of gauge” illustrates two points. Firstly, that common standards can be a huge benefit to society at large. Changing data to meet specific but arbitrarily different data standards is a waste of bandwidth. If a geospatial or Earth observation company thinks this is clever or in some way a moat, it’s not. It’s a barrier to adoption and a needless burden on those providing analytics. However, it highlights my general point about the need for robust midstream geospatial software engineering.
Secondly, seemingly generic products, such as railways designed to be fast and efficient, can be bogged down by poor engineering practices and possibly avoidable human influences, such as professional ego.
This is all to say, there are costs when we forget to think about overtly engineering the transportation layer. This is as true with data as it is with trains.
So when I talk about the midstream, what am I talking about? In short, it’s really all those activities between data capture or acquisition and product delivery. But let’s dig a little deeper.
One of the subjects I have ranted a lot about recently is trust. In the midstream, we can undertake numerous activities that might alter the value of a pixel or data point. The tracking of those activities creates a chain of custody, or the providence of the reflectance value detected by a sensor or the measurement recorded by the SAR. That providence can be used to adequately determine the fitness for purpose of a particular data product. We can know what was captured, and then what happened to a particular pixel, point or stream. Providence is closely related to metadata, and metadata management is often the basis of a data standard. Open standards, as alluded to above, are absolutely key to a common understanding of data products, but, of course, standards are always interpreted differently (ideally, not that differently, though).
Scale is also key to the midstream discussion. Scale is a loaded term in geospatial; it can mean anything from a measure of geographic relevance of a data product to the scalability of a business to the ability of a data infrastructure to serve increasingly large loads. In this case, I am referring to the last definition: serving large amounts of data to large numbers of people in a short amount of time. This is a function of midstream engineering. Whether we are leveraging cloud or cloud-like technology, the nature of the distribution infrastructure can make or break a data product's fitness-for-purpose. Latency, for instance, is the time it takes for a product to be delivered; in effect, it measures data delivery speed. For maritime domain awareness, latency is a critical metric. In fact, there are a number of mid-resolution SAR providers for whom 90% of their imagery revenue is made within an hour of collection. For other constellations, archives are more notable, so making more data available becomes the differentiator. How data is engineeringa nd delivered, therefore, defines the product suite that a data stream can support.
Data discovery used to involve calling a salesperson. Now, data discovery is more often, but not exclusively, a catalogue or an API call. This again is the midstream. Sometimes, depending on the organization in question, a phone call might be easier than using their API, but that’s not how it’s supposed to work! APIs, these days, are giving way to MCP environments, meaning AI can have a conversation with a data catalogue. This workflow is increasingly important.
Finally, once data has been discovered, it needs to be accessed, either by streaming for those more cloud-native among us, or by download, for those who need to store the pixels themselves. This data access capability can significantly affect the latency and usefulness of a data product.
So, when I talk about midstream engineering being a business driver, this is what I mean. Scale and the ability to manipulate and deliver data fit for specific purposes create new business and greater adoption. How data is delivered can determine which market segments a data product can serve. How data is engineered can have a complementary effect on an industry. By that I mean the fractional distillation of image data into a particular analytic can suddenly convert a pixel-based product, which would typically be inaccessible to many commercial applications, into a number in a database, or an indication of the presence or absence of a landscape feature. A metric which is accessible to many more customers.
The nature of midstream geospatial engineering can literally create new products. So when I say that data is not oil, yet use the term “fractional distillation”, which seems confusing, what I am pointing to is the idea that a particular set of pixels can be distilled repeatedly to create different products. That distillation is all the midstream.
To bring this discussion back to the idea of the break in gauge, I will remind you all that, over time, countries generally agreed on the gauge of their railways because the switching costs were enormously painful. So, with the standardization of railways came greater sophistication and efficiency in their use, whether for freight or people. But today, when we overlook the value of the midstream, we are offloading this cost onto customers, reducing the likelihood that they will adopt geospatial or Earth observation products.
Again, this layer has seemed invisible, yet as AI adoption spreads like wildfire, sophisticated geospatial companies will need to refine their midstream stack to keep up and engage readily.

