All work

Real-time sign recognition · 2023

Recognizing hand signs from a webcam in real time.

I built a webcam pipeline that detects a hand, converts its landmarks into normalized features, and predicts a sign label with a Random Forest classifier. The result is drawn live with OpenCV.

Role
Solo, self-directed
Stack
Python · OpenCV · MediaPipe
Classes
8 sign labels
Status
Prototype, evaluated

The problem

Recognizing a hand sign from a live camera feed requires more than a classifier. The pipeline has to find the hand, represent it consistently as the hand moves around the frame, and return a label quickly enough for the result to feel live.

What I built

The project uses MediaPipe Hands to detect one hand and return 21 landmarks. The x and y coordinates are normalized by subtracting the minimum coordinate, producing 42 features that are less sensitive to where the hand appears in the frame.

A scikit-learn Random Forest is trained on those landmark vectors. During inference, OpenCV captures the webcam frame, MediaPipe draws the hand skeleton, and the classifier's label is displayed above the detected hand.

The numbers

These are the values printed by the repository's training notebook:

Training and evaluation results
Input representation42 normalized x/y landmark features
ClassifierScikit-learn Random Forest
Classes8 sign labels
Test split20%, stratified and shuffled
Reported accuracy100.00% of held-out samples
Data cleaningInvalid 84-length samples skipped; 42-length vectors retained

The 100.00% figure is the notebook's single held-out split, not a guarantee of real-world performance. The repository does not report cross-validation, a separate test recording, confusion between similar signs, or webcam performance such as FPS and latency.

Example runs

A live run is the webcam loop in inference_classifier.py: OpenCV reads a frame, MediaPipe draws the hand skeleton, the Random Forest predicts a label, and that label is drawn above the hand.

What you see on a good frame

  • 21 MediaPipe landmarks connected as a hand skeleton
  • A bounding box around the detected hand
  • One of the trained sign labels printed above the box

There is no saved clip or confusion-matrix figure in the repository. The 100% figure above is from the training notebook's held-out split, not from a timed webcam session with measured FPS.

What went wrong

The evaluation is useful but narrow. A random split can be optimistic when samples from the same collection session appear in both training and test data. The live application also assumes one visible hand and uses a fixed label mapping, so changes in lighting, pose, background, or hand orientation may reduce reliability.

What I would do differently

I would collect separate recordings for training and testing, report a confusion matrix for each sign, and measure FPS and end-to-end latency on the target machine. I would also add an explicit unknown or low confidence state rather than forcing every detected hand into one of the eight labels.

The complete implementation is available in the Sign-Language-Detection project folder.