Digital Tower Sensor Fusion Integration Approaches
Merging radar, cameras, and ADS-B into one trusted picture.

A digital tower's value comes down to one question: how well does it merge what its sensors see into a single picture a controller can trust? No individual sensor, no matter how capable, answers that question alone.
Digital tower operational value and sensor fusion
A controller in a traditional cab works from a direct, unmediated view out the window. A digital tower has no window. It has to rebuild that view from sensor data, and the rebuilt version is only as good as the logic that stitches the pieces together. Radar, ADS-B, multilateration, cameras, and environmental sensors each cover a different slice of the operational picture, and each carries its own accuracy profile, update rate, and way of failing. Cameras lose detail in low light, fog, or when a hangar or terminal building blocks the view. ADS-B gives clean position data but only for aircraft equipped to broadcast it, and that broadcast can be spoofed or simply absent. Radar picks up aircraft that aren't cooperating with any broadcast system at all, but its update rate and resolution both get worse at close range, exactly where tower operations need the most precision. Multilateration checks position by timing when a signal reaches several ground receivers. It only works where enough receivers sit in range to begin with.
Running all four in parallel, each on its own, produces four separate pictures. None of them, taken individually, gives a controller enough to act on with confidence, and none of them, stacked side by side without a common framework, resolves disagreements between what one sensor claims and what another shows. The integration architecture makes the difference: how the system decides which input to trust, how it weighs conflicting reports, and how it arbitrates when two sensors disagree about where an aircraft actually is. A framework from WWT, built around industrial and commercial sensor fusion, states the general principle: "a single sensor modality rarely tells the complete story. The same document notes that camera-only systems "can struggle with variable lighting, partial occlusion or visual noise," and that adding depth data from LiDAR gives the system spatial grounding that makes object detection far more reliable. That principle holds in the tower context as much as anywhere else it's been tested. The fusion architecture, not any single sensor within it, makes a digital tower work.
The sensor stack a digital tower operates
Terminal and surface traffic monitoring demands precision, low latency, and a false-alarm rate close to zero, so the sensor stack a digital tower runs is built around meeting those demands, not around any general-purpose surveillance goal. The Multi-Sensor Data Processor, or MSDP, is the integration engine built for this job. A patent description of the system (US7868812B2) states that the MSDP fuses data from one or more surface movement radars, a multilateration system with integral ADS-B sensors, flight plan data, and aircraft surveillance radars to produce a combined terminal and surface traffic picture.
Each layer in that stack earns its place because it covers ground the others miss. Surface movement radar tracks aircraft and vehicles moving across taxiways and aprons, a part of the field where ADS-B coverage often thins out and where a terminal building or hangar can block a camera's line of sight. Multilateration with integral ADS-B adds a cooperative layer on top of that: aircraft broadcasting their position get cross-checked against time-difference-of-arrival fixes computed by ground stations, which tightens both accuracy and integrity beyond what either source would give on its own. Approach radar extends the picture outward into the surrounding airspace, picking up inbound and departing aircraft well before they come within range of a surface radar or a camera. A 2025 study at an airport in China found that fusing ADS-B with approach radar within tower range sharpened track position accuracy and ground-speed calculation, extended coverage across both space and time, and gave a more accurate and more reliable read on target position, track, and speed than approach radar managed on its own.
Cameras serve as the visual stand-in for the window a controller in a traditional cab would otherwise look through, and high-definition units, including night-vision and object-tracking variants, carry that load. Western Sydney International Airport will launch Digital Aerodrome Services in 2026, and it will run 25 high-resolution cameras as part of that visual layer. Weather sensors, taxiway lighting states, and equipment indicators round out the stack, feeding context that shapes how the system interprets everything else it's tracking. None of these layers stands in for the others. Each covers a portion of the picture the rest can't reach, and that division of labor raises the next question: how does an architecture combine inputs this different in kind, timing, and reliability into one coherent track?
How fusion architectures decide what the system believes
Where in the processing chain sensor inputs get combined, whether before any interpretation happens, after each stream has already formed its own independent track, or at both points, shapes how the system handles disagreement between sources and how an error in one input spreads (or doesn't) into what the controller finally sees.
Early fusion combines raw or lightly processed data from multiple sensors before any single one of them has been interpreted on its own. That gives the fusion engine the most information to work with, but it demands that every input first be aligned to a shared reference in both space and time. Getting that alignment wrong, misregistering inputs that run on different coordinate systems, different update rates, and different latency profiles, is the main obstacle in early fusion, and it can produce ghost tracks or missed detections even when every sensor involved is working exactly as designed. Early fusion performs best when the sensors feeding it are tightly coupled in time and space to begin with, and its accuracy tends to fall off as the number of input dimensions grows, a pattern documented in other multimodal fusion work where performance drops once feature dimensionality passes a manageable point.
Late fusion takes a different approach. Each sensor stream runs through its own tracking or detection process first, and only after that does the system combine the resulting track estimates, often using a Kalman filter or some form of probabilistic weighting, into one fused track. That structure tolerates sensor dropout far better: if one input degrades or fails outright, the fusion engine keeps running on whatever remains rather than needing to reprocess a combined raw feed from scratch. The cost is that correlations visible in the raw data, patterns that two sensors might reveal together before either one commits to an independent track, are gone by the time fusion happens, and no amount of weighting at that later stage can recover them.
Hybrid architectures split the difference by fusing at more than one level: early fusion within a tightly coupled cluster, such as combining ADS-B and multilateration directly, and late fusion across clusters, such as combining that already-fused cooperative track with radar and camera detections. The MSDP reflects exactly this logic. It takes in multiple radar sources alongside a combined multilateration and ADS-B feed and produces one unified surface and terminal picture, which requires cooperative surveillance, ADS-B and multilateration, and non-cooperative surveillance, radar, to get handled at different points in the pipeline. The Tianfu Airport study makes the case for why that split matters: ADS-B gives a high update rate but only works for equipped aircraft, while approach radar covers a wider area but updates less often and carries no cooperative identity. Fusing the two in a hybrid arrangement closes the gap either one leaves standing alone.
Precision, update rate, and false-alarm discipline in a fused picture
A controller working final approach or ground movement has almost no margin for position error, a delayed update, or an alert that turns out to be nothing. The proximity of aircraft to each other, the speed differentials involved, and the geometry of the airport surface all leave less room for uncertainty than an en-route controller ever has to tolerate. Monitoring an aircraft's position against the glide path or localizer, for instance, demands precision an en-route picture simply doesn't need.
Update latency has a direct, measurable cost. When the interval between surveillance updates runs long, controllers respond by applying wider speed and spacing buffers to cover the uncertainty, and that caution reduces runway throughput directly. Cutting that latency is one of the stated goals behind fusion-architecture design in tower operations, not an incidental benefit.
False-alarm rate carries equal weight as a design requirement. A system that throws frequent false alerts teaches controllers to tune them out, and once that happens, the alert loses the safety value it was built to provide. One airport study lists improved track position accuracy and ground speed calculation as direct results of its surveillance-radar fusion approach, so both count as actual operational requirements, not secondary figures nobody is held to.
Camera feeds add a further demand on top of these. Since they stand in for the window a controller used to look through directly, they need processing and presentation fast enough and clear enough that a controller can make a runway occupancy call with the same confidence direct observation would have given. These standards, precision tight enough for approach monitoring, latency low enough to protect throughput, and a false-alarm rate controllers can actually trust, set the bar any fusion architecture has to clear. A shortcut that passes muster in a warehouse or on a factory floor doesn't belong anywhere near tower operations.
Where AI processing sits in the fusion stack
AI processing sits above raw sensor fusion rather than inside it, reading patterns in the already-fused picture that no single sensor and no fixed rule set would catch reliably. Applied to high-definition camera feeds, AI image processing can flag aircraft movements, catch runway incursions, and identify weather hazards as they happen, giving controllers usable visibility in fog or low light where the raw camera feed alone would leave them guessing. Machine learning models can also watch flight paths against historical traffic and the live fused picture to flag a potential conflict before it crosses the threshold that would otherwise demand a controller's immediate action, shifting part of the job from reacting to an alert toward catching the problem on approach.
WWT's framework describes the mechanism behind that shift: pairing a computer vision model with multi-sensor object tracking lets the system run its analysis "on stable identities rather than raw detections." That distinction carries particular weight in a tower, where the same aircraft has to stay correctly identified as it moves across handoff zones between one sensor's coverage and the next.
The obvious objection holds: adding an AI layer to a safety-critical system risks adding latency or unpredictability exactly where neither is tolerable. The answer built into these architectures is structural. AI runs as an advisory and alerting layer on top of the fused picture, not as the engine generating the primary track. Surveillance stays deterministic, and the system gains an added layer of interpretation without handing track generation itself over to a less predictable process.
Fusion architecture and resilience when sensors degrade or fail
A well-built fusion architecture keeps producing a usable, if reduced, picture when a sensor input degrades or drops out. If one is built poorly, that failure passes straight through to the controller's display. Redundant sensor streams only pay off if the fusion engine can actually detect that one of them has degraded, down-weight or exclude it, and flag the reduced-confidence state to the controller. None of that happens automatically just because multiple sensors happen to be running at once. Late and hybrid architectures handle this better than early fusion on its own, because each stream forms its track independently, so one sensor's failure can't corrupt the pipeline feeding the others.
Over a two-month stretch, the airport recorded eleven runway and taxiway conflicts against a normal expectation of roughly three, and investigators have been examining fatigue, short handovers between shifts, and controllers covering combined roles as possible contributing factors. The investigating safety authority has six incidents from that period under formal investigation as of late 2026. Its preliminary investigation, running from September into October 2026, names fatigue and understaffing as areas still under examination, without yet determining whether either one actually contributed to the incidents. A continuously running, sensor-fused system doesn't get tired and doesn't need a handover between shifts, which is precisely the kind of failure mode this case puts on display.
If staffing is the binding constraint at a small or non-towered airport, a resilient fusion architecture is the entire operational case for installing a digital tower, the alternative being no air traffic control coverage. Seven airports now running one remote tower vendor's system show what that looks like in practice: real-time data integration paired with high-resolution surveillance lets a single system manage low-density traffic efficiently, in conditions that look a lot like the small U.S. airports currently operating with no tower whatsoever.
The U.S. certification gap that keeps technically sound fusion architectures from deploying in the NAS
Technical capability isn't what keeps digital tower sensor fusion out of most U.S. airports. The obstacle is regulatory: the national aviation regulator has not laid out a certification pathway for these systems, which leaves architecture that already runs successfully elsewhere legally unable to deploy at scale in domestic airspace. The FAA Reauthorization Act of 2024 required the agency to build a program and publish milestones that cover system design and operational approval for remote and digital towers. As of May 2025, that program and those milestones had not been published.
Sensor fusion has largely been solved in the places where it's been tried: approach radar and ADS-B fused together at Tianfu, surface radar and multilateration combined through the MSDP architecture, camera networks running at Western Sydney, remote systems operating across seven Norwegian airports. What's missing in that country is the regulatory structure that would let a comparable architecture carry the legal authority to run a tower on its own. Until the regulator defines that pathway, the gap between what these systems can already do and what they're permitted to do at scale in domestic airspace stays exactly where the 2024 reauthorization left it: acknowledged, mandated, and still open.


