arXiv PrePrint
Multimodal AI
BWC Analysis
September 2026
Mamadou K. Keita, Angela Srbinovska, Anita Srbinovska, Nishka Desai, Isabella Zicari, P. Kwaku Sanaah-Faried, Sanjay Charitesh Makam, Wyatt Auten, Vivek Senthil, Hannah Desnick, Jonathan Bateman, Adrian Martin, Christopher Homan, John McCluskey, Ernest Fokoué
Abstract: We introduce OmniEye, a multimodal video intelligence system for law-enforcement training and review (source code available on request to verified law-enforcement and public-safety agencies). OmniEye ingests body-worn camera footage and perceives every 30-second window jointly across video and audio with one multimodal foundation model. It then stores the model's structured output in an embedded SQLite database with BM25 full-text search. Officers can question the footage through an agent that writes structured queries, retrieves candidate windows, and re-perceives them with the model before it may cite them. The whole system runs on one 16 GB GPU with a 4-bit quantization-aware-trained model, and it also scales to full bf16 precision on a multi-GPU cluster.
arXiv PrePrint
Speech Recognition
Model Adaptation
July 2026
Vivek Senthil, Ernest Fokoué
Abstract: Modern policing faces a "visibility paradox" where law enforcement agencies possess petabytes of Body-Worn Camera (BWC) footage that remains largely unutilized for accountability or systemic review due to the prohibitive labor costs of manual transcription. This research presents a framework for adapting the OpenAI Whisper architecture to the unique acoustic and linguistic challenges of the policing environment. By employing Parameter-Efficient Fine-Tuning (PEFT) through Low-Rank Adaptation (LoRA), we address the significant performance degradation observed in zero-shot models when confronted with high-stress scenarios, sirens, and radio interference. Crucially, we demonstrate that this adaptation is feasible on consumer-grade hardware (Acer Nitro local machine with NVIDIA 4GB GTX GPU) using 8-bit quantization and gradient checkpointing. We further integrate these transcriptions into a symbolic reasoning pipeline using a domain-specific ontology to transform raw audio into evidence-linked incident graphs, achieving a 93.7% lexicon mapping rate for the advancement of procedural justice and transparency.
arXiv PrePrint
Computer Vision
BWC Analysis
May 2026
Angela Srbinovska, Christopher Homan, Adrian Martin, Ernest Fokoué
Abstract: Law enforcement agencies are accumulating vast amounts of body-worn camera (BWC) footage. However, this remains operationally opaque. That is, analysts and trainers still have to invest considerable time watching full-length videos to pinpoint the start of key encounters and identify the points where activity shifts to something more physically intense. We present an approach to process BWC video into a time-aligned sequence of fixed-length 10-second windows, processed and labeled using a privacy-conscious protocol. Each window is labeled with two dimensions of information: (i) the operational context of the window and (ii) the level of motion intensity within the window, with low-evidence labels for windows for which insufficient evidence exists due to darkness, blur or occlusion. We train models to classify windows based on these two axes using frames sampled from each window encoded using CLIP model and aggregated into a window-level representation. We extract dense optical flow statistics for each window to capture motion intensity. On test windows the best context model achieves 78.75% accuracy, and the best-accuracy activity model achieves 88.33%. We also included integrity audits to show the results and how the visual timeline representations support faster incident review and make the officer training workflow more practical.
arXiv PrePrint
NLP
Knowledge Representation
May 2026
Anita Srbinovska, Jansen Orfan, Adrian Martin, Ernest Fokoué
Abstract: Law enforcement reports contain structured fields and written narratives. However, many incident facts that are needed for review, police training, and investigations are in natural language and require manual reading. We propose a framework using symbolic methods for converting narratives into evidence-linked facts. Our objective is to measure the value of narratives to recover incident details only from the unstructured text and build temporal graphs with time cues and domain axioms. We achieve this by redacting personal identifiers, semantic parsing, predicate mapping to ontology, and reasoning. We evaluate the symbolic approach on 450 property crime reports and a short human review. Of the extracted events from the system, 54.1% had a confidence score of at least 0.80 and 93.7% were mapped through the PropBank-VerbNet-WordNet semantic path. 100% agreement was reached on incident initiation, stolen items, and temporal cues and lower agreement for forced entry interpretation.
CrimRxiv Preprint
Criminology
May 2025
Angela Srbinovska, Jonathan Bateman, Anita Srbinovska, John McCluskey, Adrian Martin, Ernest Fokoué
Abstract: Data from police body-worn cameras (BWCs), once limited to evidence collection, have evolved with the emergence of digital tools as a mechanism for near real-time analysis of police behavior and community interactions. We propose a novel interdisciplinary strategy designed to more fully leverage the capabilities of BWCs by placing equal value on the social and mathematical sciences. By fusing social science frameworks such as Systematic Social Observation (SSO) and Social Interactionist Theory (SIT) with cutting-edge artificial intelligence (AI) techniques—natural language processing (NLP), machine learning (ML), and computer vision (CV)—we offer a robust lens into police-civilian encounters that ensure accuracy, replicability, and explainability in our findings. This approach transforms how knowledge is extracted and interpreted from BWC footage, with a strong emphasis on transparency to maintain research integrity. Our fundamental goals are to detect, recognize, classify, and analyze patterns in BWC footage. This would then offer a platform for actionable recommendations for police training and building community trust and safety. We propose a methodological framework that combines advanced technological tools with contextual understanding to advance the frontiers of knowledge discovery from police BWCs.
ArXiv Preprint
AI Ethics
May 2025
Anita Srbinovska, Angela Srbinovska, Vivek Senthil, Adrian Martin, John McCluskey, Jonathan Bateman, Ernest Fokoué
Abstract: This paper proposes a novel interdisciplinary framework for analyzing police body-worn camera (BWC) footage from the Rochester Police Department (RPD) using advanced artificial intelligence (AI) and statistical machine learning (ML) techniques. Our goal is to detect, classify, and analyze patterns of interaction between police officers and civilians to identify key behavioral dynamics, such as respect, disrespect, escalation, and de-escalation. We apply multimodal data analysis by integrating image, audio, and natural language processing (NLP) techniques to extract meaningful insights from BWC footage. The framework incorporates speaker separation, transcription, and large language models (LLMs) to produce structured, interpretable summaries of police-civilian encounters. We also employ a custom evaluation pipeline to assess transcription quality and behavior detection accuracy in high-stakes, real-world policing scenarios. Our methodology, computational techniques, and findings outline a practical approach for law enforcement review, training, and accountability processes while advancing the frontiers of knowledge discovery from complex police BWC data.