The problem
Count how many people pass through a doorway or corridor over the length of a video, distinguishing people
entering from people leaving. A frame-by-frame detector alone cannot do this: it tells you a person is
present, not whether it is the same person as the frame before. Without identity, one person standing in
view gets counted once per frame.
What I built
The pipeline runs per frame. YOLOv8n detects people; a tracker assigns each detection a persistent ID
across frames; a counting region tallies an entry or an exit when a tracked ID crosses it.
- Read the frame. OpenCV pulls frames and the source width, height and frame rate.
- Detect and track.
model.track(im0, persist=True). The
persist flag is what carries IDs between calls; without it every frame starts fresh.
- Count crossings. A four-point region spanning the frame, with separate in and out
tallies.
- Write the output. Annotated frames with boxes, IDs and trails written back to a video
file.
The counting region is a shallow band across the full frame width rather than a single line:
region_points = [(20, 400), (1080, 404), (1080, 360), (20, 360)]
A band gives the tracker a few frames to register the crossing. A one-pixel line can be stepped over
between frames by anyone moving quickly, and the crossing is missed.
The numbers
Pipeline configuration
| Detector | YOLOv8n, pretrained COCO weights |
| Tracking | Persistent IDs, trails drawn for debugging |
| Counting region | Four-point band, 1060px wide, 44px deep |
| Input | Recorded 1080-wide video |
| Attribute model | DeepFace, ethnicity action only |
| Model evaluation | Pretrained models; no project-level benchmark |
The repository uses pretrained YOLOv8n and DeepFace models, but it does not include a labelled test set or
a project-specific accuracy report. I watched the annotated output and the numbers looked right. That is
not the same as measuring whether the system is correct.
The half I argued against
The second requirement was ethnicity estimation per detected person. I implemented it with DeepFace,
calling the ethnicity classifier on each detected face, and I do not think it should be deployed.
The call runs with enforce_detection=False. That flag exists so the pipeline does not crash
when no clear face is found, but it means the classifier returns a confident label regardless. A blurred
face, a back of a head, a person at the edge of the frame: all get a label, all get a confidence score,
none of it is grounded in a face the model actually resolved.
Underneath that, the model predicts a categorical ethnicity label from appearance. Its accuracy is not evenly
distributed across the groups it claims to distinguish, its categories are the ones its training set
happened to use, and there is no threshold at which a wrong label here is harmless. Counting footfall is a
reasonable thing to automate. Sorting people into racial categories from CCTV is not the same kind of task,
and treating it as one more model call is the error.
I built it because it was the assignment, and I wrote up why it should not be used. Given the choice now I
would propose a different second half: dwell time, queue length, or occupancy against a capacity limit, all
of which answer the underlying operational question without classifying anybody.
Example runs
The repository contains annotated output videos under
Test and Results
(human_counting_video.avi, combined_video.avi). I am not embedding stills here:
the footage shows identifiable faces, and this page already discusses an ethnicity classifier.
What a counting run looks like
On a recorded 1080-wide clip the pipeline writes an annotated video with:
- YOLOv8 person boxes and persistent track IDs
- Short motion trails used while debugging the tracker
- A shallow counting band across the frame (
region_points = [(20, 400), (1080, 404), (1080, 360), (20, 360)])
- Separate running tallies for entries and exits as IDs cross that band
That is a demo of the overlay, not a measured accuracy result. No hand-labelled minute of footage was
scored against these counts.
What I would do differently
Hand-label one minute of footage and measure counting accuracy against it. Without that number the project
has no result, only a demo.
Drop the attribute stage and replace it with something that answers the operational question without
classifying people.
Test the counting band against people moving at different speeds. I tuned those coordinates by watching one
video until the numbers looked right, which means they are fitted to that video and I do not know how they
behave on another.