
In university lecture halls and corporate environments with over 60 students per session, taking roll call manually consumes 10 to 15 minutes of valuable class time. Beyond the productivity loss, paper sign-in sheets and basic card swipe terminals are vulnerable to proxy attendance, physical credential sharing, and data entry discrepancies.
During my Bachelor of Engineering in AI & Machine Learning at Chandigarh University, my team and I set out to engineer an automated, touchless attendance verification system powered by real-time computer vision and edge facial recognition.
Our goal was to design a production-grade edge application capable of detecting multiple faces simultaneously from standard camera feeds, extracting discriminative facial embeddings, matching identities against a student database in sub-50ms, and updating attendance logs in the cloud with zero manual intervention.
The Bottlenecks of Traditional Attendance Systems
Attendance tracking seems deceptively straightforward until scaled across an academic institution with hundreds of classrooms and thousands of students. Conventional workflows suffer from three primary architectural failures:
1. Latency & Disruption: Calling out names or passing around paper sheets disrupts the lecture flow and takes up to 20% of contact teaching time.
2. Proxy Marking: Manual systems have no cryptographic or biometric verification. A student can sign on behalf of an absent peer without consequence.
3. Database Latency: Paper sheets must be manually tabulated into central enterprise resource planning (ERP) databases days later, eliminating any opportunity for real-time absence notifications or early interventions.
Biometric facial recognition solves these issues simultaneously: it is contactless, inherently identity-bound, and can be automated to log attendance the instant a student enters the classroom.
Computer Vision Pipeline: Detection, Alignment & 128-D Encodings
Our recognition pipeline is structured into three discrete stages: face detection, landmark affine normalization, and deep metric embedding projection.
First, each incoming 1080p frame from the video stream is downsampled to 1/4 resolution to ensure 30+ FPS throughput. Faces are detected using an optimized Histogram of Oriented Gradients (HOG) detector paired with a linear SVM, providing high precision on frontal and semi-profile orientations without requiring heavy GPU compute.
Once candidate bounding boxes are located, a 68-point facial landmark detector identifies key physiological anchors: eyes, nose bridge, nostrils, and chin perimeter. The face crop is warped via an affine transformation to normalize inter-pupillary distance and tilt angles.
The aligned face crop is passed through a deep residual network (ResNet) trained with triplet loss to map the facial features into a compact 128-dimensional Euclidean space. In this metric space, encodings of the same individual are clustered closely together, while encodings of different individuals are separated by a minimum Euclidean margin.
import cv2
import face_recognition
import numpy as np
from datetime import datetime
class EdgeAttendanceEngine:
"""Real-time multi-face recognition with debounce cooling off."""
def __init__(self, known_encodings, student_ids, threshold=0.52):
self.known_encodings = known_encodings
self.student_ids = student_ids
self.threshold = threshold
self.attendance_log = {} # {student_id: timestamp}
def process_frame(self, frame, debounce_seconds=300):
# Downsample for sub-30ms real-time throughput
small_frame = cv2.resize(frame, (0, 0), fx=0.25, fy=0.25)
rgb_frame = cv2.cvtColor(small_frame, cv2.COLOR_BGR2RGB)
locations = face_recognition.face_locations(rgb_frame)
encodings = face_recognition.face_encodings(rgb_frame, locations)
marked_students = []
for encoding in encodings:
distances = face_recognition.face_distance(self.known_encodings, encoding)
best_idx = np.argmin(distances)
if distances[best_idx] <= self.threshold:
student_id = self.student_ids[best_idx]
now = datetime.now()
# Anti-proxy debounce: prevent redundant writes
last_time = self.attendance_log.get(student_id)
if not last_time or (now - last_time).total_seconds() > debounce_seconds:
self.attendance_log[student_id] = now
marked_students.append((student_id, distances[best_idx]))
return marked_studentsReal-Time Synchronization & Anti-Proxy Throttling
Identifying a face is only half the engineering challenge; persisting and synchronizing verified attendance records without database lockups is equally critical.
We integrated Firebase Realtime Database and Cloud Storage. Student registration images are pre-processed and stored in Firebase Storage, while their 128-dimensional vector signatures are cached in local memory on the edge device during system initialization.
When a verified match is confirmed, the system immediately writes to the date-partitioned JSON tree (`/Attendance/{YYYY-MM-DD}/{student_id}`). We implemented an anti-proxy debounce threshold: once a student is logged, their ID enters a 5-minute cooling-off window. This prevents thousands of redundant cloud writes as students sit facing the classroom camera.
Edge Performance Under Variable Lighting & Angles
Classroom lighting is notoriously unpredictable: bright morning sunlight, flickering fluorescent tubes, and backlight from open windows can distort raw pixel intensities. To guarantee high recognition fidelity, we applied Contrast Limited Adaptive Histogram Equalization (CLAHE) on the luminance channel prior to feature extraction.
In empirical testing across 120 students, the system achieved a 96.8% verification accuracy with a false acceptance rate (FAR) below 0.05% when the Euclidean distance threshold was tuned to 0.52. End-to-end recognition and database persistence completed in under 42ms per frame on a standard Intel Core i5 edge device.
“Biometric edge systems succeed or fail not on raw classification accuracy, but on latency, debounce management, and graceful handling of variable real-world lighting.”
Key Architectural Lessons from Campus Deployment
Deploying this system in active academic environments taught us critical real-world systems lessons:
First, offline-first edge architecture is non-negotiable. If campus Wi-Fi drops momentarily, the recognition pipeline must continue running locally, buffering attendance events in a local SQLite ring buffer and synchronizing to Firebase once connectivity resumes.
Second, user feedback must be instant and ambient. We built an OpenCV HUD overlay that outlines recognized students in green with their name and student ID while displaying attendance confirmation for 2 seconds. This visual feedback gives students immediate confidence that their presence was registered without needing to stop or check an app.
The project demonstrated that combining lightweight computer vision pipelines with modern serverless cloud databases can completely eliminate attendance friction in institutional environments.