INSIGHT

The intelligence layer for physical security: why cameras still can't tell a threat from a false alarm

There are a billion CCTV cameras in the world and almost no understanding of what they see. Why the physical security market's next decade is about judgement, not capture.

16 July 2026

The world has never had more cameras, and they have never been less watched.

There are around a billion CCTV cameras deployed globally. On any given night the overwhelming majority are recording faithfully to a hard drive no one will ever open, unless something goes wrong badly enough to justify trawling back through the footage. We have spent three decades installing the eyes. We have barely begun to build anything that can make sense of what they see.

This is the strange shape of the physical security market in 2026. The hardware problem is solved. The understanding problem is not.

5 million commercial CCTV cameras in the UK

How a monitored site actually works today

It is worth being precise about how monitoring works today, because the gap hides in the detail. On a typical site the cameras record continuously. Motion triggers an alert. A layer of AI strips out the obvious noise. And then, thirty to ninety seconds after the event, a human reviews a ten to fifteen second clip and decides whether it matters.

Look at that chain and you notice something. Every step captures, records or filters. Not one of them judges. The judgement - is this a threat, and who needs to act - still lands on a person. Often late, because the review comes after the event. Often not at all, because one operator watching a wall of screens across a dozen sites cannot be everywhere at once.

The industry's answer to this has been better filtering. Motion alerts fire far too often - on foxes, on shadows, on weather, on a tarpaulin flapping in the wind - so a layer of AI strips the obvious false positives before they reach a human. This genuinely helps. It is also the wrong tool for the real problem, because a filter that can tell a person from a fox still cannot tell a gardener from an intruder. Everything that requires actual understanding still lands on the human, if it lands at all.

The predictable result is alert fatigue. A system that cries wolf a hundred times a day trains the people watching it to stop listening. We came across a site that had spent tens of thousands on a new camera system and then switched the alerts off within weeks, because the noise was unbearable. A break-in went unnoticed until the morning. That is not a story about bad cameras. It is a story about cameras that capture everything and understand nothing.

The thing that changed

For most of CCTV's history, the technology to close this gap did not exist. Machines could detect objects - a person, a vehicle, a bag left on a platform - but detecting an object is not the same as understanding a scene. Object detection tells you what is in the frame. It cannot tell you what is happening, and "what is happening" is the only question a security operator actually cares about.

Vision language models changed that, and only recently. For the first time, a system can read behaviour and context rather than mere presence. Not "person detected" but "someone is climbing the perimeter fence rather than using the gate, wearing no hi-vis, in a way consistent with an intrusion". The distance between those two sentences is the distance between an alarm and an answer. One makes a tired operator look up. The other tells them what to do.

A category being built in real time

None of this is a secret, and some of the most credible investors in the world have noticed. Over the past eighteen months, companies including Augur, Hakimo, Coram and Lumana have raised well over $150m between them to build exactly this understanding layer on top of existing cameras. We read that as validation, not as a threat. The question is no longer whether this layer gets built. It is which architecture, and which route to market, wins each part of the market.

Because the approaches genuinely differ, and the differences matter. Some run their inference in US cloud, which sends site footage across the Atlantic. Some sell direct to large US enterprises. At least one is building top-down for governments and critical national infrastructure, with the long procurement cycles that implies. These are not cosmetic distinctions. They decide where a site's data lives, how fast an alert arrives, whether the system keeps working on a poorly connected rural compound, and who owns the relationship with the customer. The layer is agreed on. How to build and deliver it is very much not.

CCTV camera mounted on a pole overlooking a site

Where we think the answer lies

Our own view - and it is a view, held in a contested market, not a claim that others cannot build - is that three choices matter most for the commercial mid-market. The first is edge-first architecture: running detection on the site itself, so the system keeps working when the connection does not, and so a site's footage stays where it should. The second is data sovereignty, which for UK commercial clients with real processing obligations is a contractual requirement, not a nice-to-have. The third is the channel: the security firms and installer-monitors who already install, maintain and monitor these cameras, and who already own the customer relationship and the billing.

We should be honest about where the architecture stands today. Full edge inference is the direction we are heading; at present detection runs at the edge and the heavier reasoning runs in UK-only cloud. We do no facial recognition - the analysis is behavioural, which for this market is a feature rather than a limitation.

And we would argue the durable advantage here is not the model. Models commoditise; this year's frontier is next year's baseline. The advantage is the labelled, operator-checked outcome data and the per-site context that accumulate with every deployment - a corpus that a professional has already verified, rather than a pile of unlabelled footage. That is the thing that gets sharper the more the system is used, and it is not something a well-funded newcomer can buy overnight.

The point

The market does not need a better filter. It has plenty of those. It needs the system to understand the site - to know the difference between a delivery and a break-in, between a contractor and a trespasser, between the ten seconds that matter and the ten thousand that do not.

The cameras are already on. The whole of the next decade in physical security comes down to a single question: is anything finally watching?

The intelligence layer for physical security

We work with UK security firms, warehouses, construction sites and commercial property operators, turning the cameras already on site into a real-time watch that understands what it sees.

Book a Demo