A camera sees pixels. It does not see intent, context, or consequence. Modern detection can tell you a great deal about what is in the frame, and nothing at all about whether it matters.
Understanding that boundary is the difference between a system that floods you with alerts and one that tells you the things you actually needed to know.
Detection raises a candidate
Good detection is genuinely impressive. It can flag a person where there should be none, a vehicle after hours, movement along a fence line. What it produces is a candidate: something worth a second look.
A candidate is not a conclusion. Treating every candidate as an incident is how alert fatigue starts.
Context is where it stops
Whether the person by the fence is a threat or a contractor, whether the after-hours vehicle is a problem or the owner, is a question about context the camera does not have.
Software can narrow the field remarkably well. It cannot close it, because the last step is judgement, not detection.
Where the operator begins
This is the point at which a person takes over. A trained operator looks at the candidate the system raised and answers the question the camera cannot: is this real, and does it need a response.
The machine does the watching that people are bad at. The person does the deciding that machines are bad at.
Built around the boundary
Ocular is designed around exactly this line. Detection runs continuously and surfaces candidates; an operator verifies before anything escalates. Neither is asked to do the other’s job.
A camera cannot tell you what matters on its own. It was never supposed to. The system around it is what makes the footage mean something.