Products
Measuring indoor air quality (IAQ) shows us how a space is doing right now. Once a high CO₂ peak has occurred, however, the decision comes too late. The occupants have already been exposed, and ventilating all at once may carry an energy cost that could have been avoided.
Hence the interest in anticipating. If we can know that a room will reach elevated levels two hours from now, the conversation changes entirely. Ventilation can be adjusted ahead of the peak, the load on the HVAC system can be spread out over time, and the space can be kept in healthy conditions without over-ventilating.
The difficulty is that, until recently, the scientific literature offered very little methodological guidance for tackling this in an IAQ context (most studies to date have focused primarily on outdoor air quality). Many were confined to a single building or to short monitoring periods, or compared models without consistent criteria. And hardly any answered the most practical question of all: which CO₂ forecasting methodology should I actually use?
In an initial study published in the journal Sensors, we addressed precisely that gap, focusing on a particularly sensitive environment: school classrooms. CO₂ was monitored in fifteen schools across the Pamplona metropolitan area using MICA devices, and a range of forecasting strategies was compared across several time horizons, from 10 minutes to 4 hours.
The most interesting conclusion was not "this model is the best", but something rather less intuitive: the answer depends on how far ahead you want to look. Over very short horizons, the simplest models are hard to beat. CO₂ evolves gradually, so assuming that levels ten minutes from now will be roughly what they are today is, quite simply, a good forecast. As the horizon lengthens, that logic breaks down. This is where machine learning models begin to earn their place, because they learn the patterns in the series itself and their accuracy degrades far more gently. The problem becomes more complex, and machine learning is able to realise its potential by learning more intricate patterns.
And it is precisely the longer horizons that carry operational value: the sooner you know what is coming, the more room you have to act. In short, forecasting is not a problem with one single solution; it is a problem that depends on the reaction time you need.
A classroom is a relatively orderly environment: fixed timetables, stable groups, a recognisable weekly pattern. The reasonable doubt was whether those conclusions hold up outside that setting.
Because indoor CO₂ is driven above all by occupancy. How many people are present, when they arrive, how long they stay, and how the space is ventilated while they are inside. And that changes radically depending on what the space is used for. A hospital bears no resemblance to a meeting room, nor an office to a food market.
Answering that question is exactly what we set out to do in the study published in the journal Forecasting, which the journal has featured in the main banner of its homepage in recognition of the attention it is receiving. We analysed 18 consecutive weeks of real-world data from 16 MICA devices deployed across six types of space, in Norway and Barcelona, within the framework of the European K-HEALTHinAIR project.

For forecasting purposes, those six use cases fall into three markedly different ways of behaving. There are spaces with intermittent, unscheduled occupancy, such as meeting and council chambers or student residences: long stretches at background levels punctuated by episodes that follow no recognisable pattern, whether sharp peaks that appear and vanish within minutes or slow build-ups that depend on whether someone opens a window. There are spaces with cyclical, recognisable occupancy, such as offices, dining halls and food markets, where a daily and weekly rhythm repeats itself —lunchtime, working hours, market day on Saturday— even if the intensity varies from one day to the next. And there are spaces with continuous occupancy, such as hospitals, with high, sustained concentrations around the clock, where mechanical ventilation matters more than the people themselves: in the hospital we analysed, CO₂ rose overnight, when airflow was reduced, and fell during the day.
Three very different dynamics, and no reason at the outset to expect a single model to perform equally well across all three.
When eleven different models are compared across these six scenarios, the pattern that emerges is clear and, in practical terms, genuinely useful: the best strategy depends on the type of space.
In spaces with sporadic occupancy (meeting rooms, residences), simple models hold up well and even beat far more complex alternatives. Deploying sophisticated architectures there means spending computational resources for nothing in return. These spaces call for a dedicated line of research addressing forecasting where occupancy is intermittent.
In spaces with cyclical patterns (offices, dining halls, markets), the story is the opposite. There is enough temporal structure to learn from, and machine learning —particularly the combinations of several models known as ensembles— consistently improves on the simpler methods at horizons of 2 and 4 hours.
The hospital proved to be the hardest case for every model without exception. Continuous occupancy and mechanical ventilation mean that the CO₂ history alone is not enough. Here it would be necessary to bring in external information, such as HVAC schedules, occupancy counts or outdoor conditions that can give the model more to work with, since the CO₂ time series on its own is not representative.
What is clear is that when occupancy generates complex but recognisable patterns, machine learning is a good solution. But the specific model changes from one use case to another. Before deploying a predictive solution, you have to look at how the space behaves.
This is where the study moves into genuinely new territory.
Recent years have seen the emergence of what are known as foundation models. Models trained on enormous corpora of data that, rather than learning one specific task, learn to generalise. They are close cousins of the LLMs now used to write code or build agents; they share the same architecture, but in this case they are pre-trained on millions of numerical time series rather than on text.
The idea is a powerful one: could a model that has never seen your building forecast the CO₂ in your building?
The results suggest that it can, on two levels:
This second point is the most significant from an operational standpoint. A newly installed sensor has no history, and until now that meant waiting weeks or months before anything could be forecast. With a foundation model, a competent forecast is available from day one, with a move to a locally fine-tuned model later on once sufficient data has accumulated.
For deploying forecasting in real-world settings, no longer depending on months of historical data for each new space changes the rules of the game.
Monitoring told us how the air is. Forecasting tells us where it is heading. Moving from one to the other is what turns air quality data into a management tool rather than just an indicator.
These two studies contribute something concrete to that journey. A reproducible methodological framework, evidence on which conclusions can be extrapolated from one type of space to another and which cannot, and a first serious evaluation of foundation models in this field. They also leave the following questions open, and we are already working on them: can a model trained in one office serve another office? Or even a dining hall with similar patterns? If the answer is yes, the historical-data bottleneck disappears.
In the meantime, the most honest conclusion of the work is also the most useful: there are no universal shortcuts, but there is a clear basis for choosing well. Translated into practical decisions, all of this comes down to: