The problem
Recognizing a hand sign from a live camera feed requires more than a classifier. The pipeline has to find
the hand, represent it consistently as the hand moves around the frame, and return a label quickly enough
for the result to feel live.
What I built
The project uses MediaPipe Hands to detect one hand and return 21 landmarks. The x and y coordinates are
normalized by subtracting the minimum coordinate, producing 42 features that are less sensitive to where the
hand appears in the frame.
A scikit-learn Random Forest is trained on those landmark vectors. During inference, OpenCV captures the
webcam frame, MediaPipe draws the hand skeleton, and the classifier's label is displayed above the detected
hand.
The numbers
These are the values printed by the repository's training notebook:
Training and evaluation results
| Input representation | 42 normalized x/y landmark features |
| Classifier | Scikit-learn Random Forest |
| Classes | 8 sign labels |
| Test split | 20%, stratified and shuffled |
| Reported accuracy | 100.00% of held-out samples |
| Data cleaning | Invalid 84-length samples skipped; 42-length vectors retained |
The 100.00% figure is the notebook's single held-out split, not a guarantee of real-world performance. The
repository does not report cross-validation, a separate test recording, confusion between similar signs, or
webcam performance such as FPS and latency.
Example runs
A live run is the webcam loop in
inference_classifier.py:
OpenCV reads a frame, MediaPipe draws the hand skeleton, the Random Forest predicts a label, and that label
is drawn above the hand.
What you see on a good frame
- 21 MediaPipe landmarks connected as a hand skeleton
- A bounding box around the detected hand
- One of the trained sign labels printed above the box
There is no saved clip or confusion-matrix figure in the repository. The 100% figure above is from the
training notebook's held-out split, not from a timed webcam session with measured FPS.
What went wrong
The evaluation is useful but narrow. A random split can be optimistic when samples from the same collection
session appear in both training and test data. The live application also assumes one visible hand and uses
a fixed label mapping, so changes in lighting, pose, background, or hand orientation may reduce reliability.
What I would do differently
I would collect separate recordings for training and testing, report a confusion matrix for each sign, and
measure FPS and end-to-end latency on the target machine. I would also add an explicit unknown or low
confidence state rather than forcing every detected hand into one of the eight labels.
The complete implementation is available in the
Sign-Language-Detection project folder.