Immersive Audio 2025: Controlling & Monitoring Immersive Audio
Object Based Audio brings an expansion in required processing capacity, increased complexity for the mix engineer, and a much bigger monitoring challenge.
All 6 articles in this series are now available in our free eBook ‘Immersive Audio 2026 – The Book’ – download it HERE.
All articles are also available individually:
All The Help We Can Get
Mixing for television used to be a relatively simple affair. Not that long ago, audio mix engineers pretty much knew what their customers were listening on, and a mono and stereo feed usually covered all the bases and could be efficiently monitored.
Not so today. Television is no longer linear, and it is no longer just television.
Since the 2000s there has been a steady growth in the number of multi-channel formats that content providers are expected to deliver. First there was the introduction of 5.1 surround which promised a more visceral experience by putting the viewer in the center of everything, and that was followed by even more immersive formats like Dolby Atmos that brought the viewer even closer to the action.
We can all get behind that. For live sports it’s a no-brainer, but spatial audio isn’t always appropriate. In some cases, like in a late-night news report, it would be downright confusing, and if it’s not adding any value, then what’s the point?
But it’s a different story when it comes to managing audio objects for personalized listening. Working with objects provides A1s with brand new creative opportunities for live sports and entertainment presentations, such as adding more crowd ambience in the height channels or immersive stings adding height and space to replay graphics and in-game statistics. But as we have discussed, objects can also make all programming more accessible, and they can help neurodivergent viewers focus on what is important.
In fact, the inclusion of these objects makes the potential number of presentations almost limitless.
All these things need to be considered and planned for, not to mention quality controlled for aesthetic and legal purposes. If we’re being brutally honest, Next Generation Audio (NGA) content like immersive and personalized audio might be an absolute boon for the end user, but it’s not making anyone’s job any easier at the other end of the chain.
Thankfully, technology vendors are helping pick up the slack.
Capacity, Control & Monitoring
The most immediate challenges are with capacity, control and monitoring.
Fundamentally, immersive and personalized audio production needs more resources. Way more. On a most basic level, it’s simple math: if a mono source uses the resources of one channel, stereo uses two, 5.1 Surround uses six and a full 7.1.4 Dolby Atmos source uses 12 to incorporate the additional side and height channels. That’s on every immersive input channel, and on every output channel.
The good news is that for the broadcast console that sits at the center of all this, this is no longer the concern that it once was.
Since the development of digital broadcast consoles in the early 2000s, manufacturers have introduced cumulative upgrades that have simplified installation and developed ever-more powerful DSP engines to keep up with the demands of broadcasters looking to produce more engaging content. And because DSP is endlessly adaptable and its firmware can be programmed to do very specific jobs, console manufacturers have been able to keep pace by developing new features and building more value into incumbent products in a relatively short space of time.
More recently, the development of true remote and distributed production workflows has further eased the burden of oversubscribed processing capacity, while cloud compute and virtualized resources deliver even greater potential for short-term boosts. Broadcasters running out of processing channels on their on-premise DSP hardware rack can now add extra capacity from another connected DSP resource on its network, or buy temporary licenses to add more from a virtualized source.
More Control
In short, capacity is no longer a barrier to generating channel-based immersive mixes from the console, and console manufacturers have also upped their game when it comes to controlling all these extra resources. Broadcast consoles all have bus width and monitoring path capabilities to cope with 7.1.4 audio as a minimum, as well as incorporating 3D panning on individual channels and balancing the control of incoming stems with height and scene-based formats. Channel-based audio can be sent straight to the encoder and defined as such in the metadata, and scene-based audio also has relative position recorded in the metadata.
When it comes to defining, creating and managing objects for personalization, most objects are fairly standard; in something like live sports these might be dialog channels for commentary, crowd atmos, or game FX.
But they still need to have clean signals, and monitoring and quality control are just as fundamental. Monitoring and QC are more than making sure things sound good and guaranteeing quality of service to the consumer. In many regions there are strict broadcast standards that content providers must adhere to, such as international loudness standards from the likes of the EBU (European Broadcasting Union), the ATSC (Advanced Television Systems Committee) in North America, and ARIB (Association of Radio Industries and Businesses) in Japan.
Loudness regulations are required to meet specific levels over the duration of a program and broadcasters have long been adhering to these with their channel-based outputs. The same applies to personalized object-based output, and having no control over the final product puts even more pressure on the audio mixers.
Upping The Ante
Here too manufacturers are adapting their services to provide more assistance to operators. There is already a fine history of audio tools that automate some of the more mundane tasks in the broadcast workflow; for example, automixers that automatically control the gain of multiple microphones in real time to minimize spill in chat show environments have been easing the pressure on beleaguered operators for decades.
Could AI up the ante? Research from the UK’s University of Salford suggests it can. Funded by Innovate UK, the AQUA (Automated Quality Control) project kicked off in 2025 with the aim to develop AI-driven automated software to address the lack of automated processes for audio QC. This adoption of AI not only seeks to automate existing processes but aims to go above and beyond basic attributes like loudness and silence detection to develop a more rounded treatment for live production and distribution, in both on-prem and cloud-based deployments.
Monitoring
When it comes to dealing with spatial objects, even more help is available. Compared to stereo listening environments, immersive monitoring is much more complex, demanding more space, more equipment, and rigorous calibration. In a cramped outside broadcast environment, reliable spatial monitoring can be difficult to achieve due to space restrictions; in fact, this is one aspect of remote production that has proved to be a real benefit as A1s can mix in environments that have been custom designed for immersive monitoring.
But the desire for more dedicated immersive environments extends well beyond broadcast. More and more recording and post houses are updating to Atmos environments and many recordings – especially classical music – are being remixed into Atmos presentations for consumers to enjoy at home and in binaural formats on headphones.
Genelec and Neumann are two manufacturers who have developed their own ecosystems to bridge the gap between professional in-room loudspeaker setups and personal headphone monitoring. Both Genelec’s UNIO System and Neumann’s RIME attempt to solve the typical HTRF challenges we addressed in article two of this guide about how our physicality filters sounds differently from person to person.
The UNIO Personal Reference Monitoring system from Genelec combines the 9320A SAM Reference Controller, Reference Measurement Microphone, and 8550A Professional Reference Headphones. Paired with the headphone calibration features in Genelec’s GLM software and Aural ID binaural headphone monitoring technology, the company says that this combination of technologies is able to reliably translate mixes between in-room loudspeaker environments and headphones for mixing on the move.
Meanwhile, Neumann’s Reference Immersive Monitoring Environment (RIME for short) aims to do the same, albeit with a slightly different approach. Neumann is part of the Sennheiser group, and Sennheiser has been working with virtual spatializers since its acquisition of the Dear Reality plugin in 2019 when the technology was added to its AMBEO Immersive Audio department.
Dear Reality was retired in 2025, and RIME was launched in April that same year. Rather than present a simulation of a space, RIME adopts a professional control room setup captured using Neumann’s KH line monitors and its iconic KU 100 binaural head-microphone. Thereby aiming to present the environment of a real optimized control room, the plug-in is built for Neumann’s NDH 20/30 headphones which the company says delivers consistent, accurate headphone monitoring. Like Genelec’s approach, it also supports a way to adjust interaural time differences to calibrate for the user’s head size and physicality.
Changing Roles
Audio production is evolving to meet the demands of NGA content, and the challenges facing mixers and broadcasters are growing in line with these demands. With such a complicated mix of objects, formats, and delivery options, there’s much more to do, but thankfully there are tools being developed every day, from ever-expanding DSP capabilities and cloud-based workflows to slimline monitoring environments and AI-driven QC solutions.
As the industry adapts to more and more object-based workflows where the consumer experience can be a variety of presentations, continued developments may even change the role of an audio mixer into one who simply creates clean, controlled audio stems that are automixed as separate objects further down the chain.
Time will tell. At the moment, and whatever all this might sound like in the future, for the time being it’s still an A1’s ears that are shaping the consumer experience. And long may it continue.
Supported by
You might also like...
The Changing Face Of Live Sports: Part 1 - The Rise Of Nimble Production
Live sports broadcasting has always been the preserve of big leagues and big broadcasters with the infrastructure, the clout and the resources to match. But it is no longer the only game in town.
Virtual Production For Broadcast: The Magic Of Subframe Technology
Subframe technology solves a fundamental challenge in virtual production, enabling LED video walls to display different images to multiple cameras at the same time. This allows each camera in a multi-camera setup to see its own perspective-correct background on a…
Standards: Audio - High Efficiency Audio Codecs (HE-AAC)
HE-AAC builds on the foundations of AAC to deliver near CD-quality audio at bitrates as low as 32 kbps, making it the codec of choice for mobile TV, digital radio and low-bandwidth streaming. This guide unpacks the key technologies behind its…
IP Security For Broadcasters 2026 – The Psychology Of Security
As engineers and technologists, it’s easy to become bogged down in the technical solutions that maintain high levels of computer security. But as the boundaries between traditional broadcast engineering and IT continue to dissolve, the first port of call i…
Standards: Audio - Advanced Audio Coding (AAC)
AAC succeeded MP3 by delivering better quality at lower bitrates. This guide examines how it works, compares the leading encoder implementations, and explains where it sits within the broader MPEG audio standards landscape.