Go to main content

Artificial Intelligence (AI) has rapidly expanded the capabilities of video understanding, yet real deployments must contend with scale, latency, privacy, and constantly evolving human behavior. This dissertation addresses these challenges end-to-end. From system architecture to learning algorithms and datasets. The effort has been made to make intelligent video surveillance both effective and responsible in the real world. The first study, Ancilia: Scalable Intelligent Video Surveillance for The Artificial Intelligence of Things, introduces an end-to-end scalable intelligent video surveillance system for the Artificial Intelligence of Things. Through empirical evaluation, Ancilia has demonstrated its ability to bring state-of-the-art artificial intelligence to real-world surveillance applications. Ancilia performs high-level cognitive tasks (i.e. action recognition and anomaly detection) in real-time, all while respecting ethical and privacy concerns common to surveillance applications. The second study is dedicated to Towards Adaptive Human-centric Video Anomaly Detection: A Comprehensive Framework and A New Benchmark. Human-centric Video Anomaly Detection (VAD) aims to identify human behaviors that deviate from normal. At its core, human-centric VAD faces substantial challenges, such as the complexity of diverse human behaviors, the rarity of anomalies, and ethical constraints. These challenges limit access to high-quality datasets and highlight the need for a dataset and framework supporting continual learning. Moving towards adaptive human-centric VAD, we introduce the HuVAD (Human-centric privacy-enhanced Video Anomaly Detection) dataset and a novel Unsupervised Continual Anomaly Learning (UCAL) framework. UCAL enables incremental learning, allowing models to adapt over time, bridging traditional training and real-world deployment. HuVAD prioritizes privacy by providing de-identified annotations and includes seven indoor/outdoor scenes, offering over 5x more pose-annotated frames than previous datasets. Our standard and continual benchmarks, utilize a comprehensive set of metrics, demonstrating that UCAL-enhanced models achieve superior performance in 83.33% of cases, setting a new state-of-the-art (SOTA).. The third study presents ARAL RGB+D Dataset: A New Action Recognition Benchmark Dataset for Surveillance Applications. Research on action recognition has advanced quickly, yet common benchmarks still underrepresent surveillance conditions where cameras differ in height and angle, and subjects are often distant or viewed from the back. In this work we introduce ARAL RGB+D, a large-scale surveillance oriented dataset with 139,171 videos and 35,181,725 frames across 60 actions, captured by seven time-synchronized cameras at 1920x1080 and 60 fps. Cameras provide frontal, angled, high-angle, and back views for every instance. Each sample includes RGB video, depth, human tracks with identities, and 2D keypoints. To investigate the factors that matter in deployment, we define three intra-dataset protocols that isolate identity (cross-subject), background (cross-setup), and viewpoint (cross-angle), and a symmetric cross-dataset transfer. By benchmarking State-Of-The-Art (SOTA) pixel and pose methods, we quantify generalization gaps and set clear targets for improvement. ARAL, not only provides a large corpus of data, but also provides a practical testbed for studying and learning out-of-domain robustness in surveillance settings. ARAL will be publicly available to the community. Together, these studies advance intelligent video surveillance along three complementary axes: scalable system design, adaptive algorithm under privacy constraints, and rigorous, surveillance-specific datasets at scale. The resulting artifacts, namely Ancilia, the HuVAD dataset and UCAL framework, and the ARAL RGB+D dataset demonstrate state-of-the-art accuracy, real-time performance, and improved generalization, while foregrounding ethical considerations. By narrowing the gap between laboratory benchmarks and operational deployments, this dissertation provides practical tools and public resources that can accelerate research and support responsible adoption in safety-critical environments. This dissertation is structured in accordance with the university’s three-paper policy. Under this framework, the dissertation comprises three independent but related research papers that collectively address the overarching research theme. Each paper contributes distinct theoretical, methodological, and empirical insights, while together they form a coherent and comprehensive body of scholarly work consistent with the university’s doctoral research standards.

Metric
From
To
Interval
Export
Download Full History