The past decade has seen explosive growth in video data (live and on-demand), driven by ubiquitous cameras, mobile devices, and online platforms. From city CCTV and body-cams to retail, sports, tele-health, and social live commerce, organizations now operate end-to-end video streaming pipelines at massive scale. In parallel, advances in deep learning, spanning video transformers, self-supervised pretraining, and multimodal large models, have transformed what intelligent systems can infer from continuous visual data, moving beyond static clips to long-horizon, low-latency understanding of events, activities, and interactions.
Affective computing adds a new dimension by interpreting emotional and social signals in video and multimedia. For example, research shows that portraying positive emotions and trust cues in video significantly increases viewer engagement. By combining facial expression, gesture, and acoustic analysis, advanced AI can infer user sentiment and intent. These capabilities enrich social-media intelligence, recommender systems, and customer journey analysis.
This special issue seeks cutting-edge work at the intersection of video analytics and affective computing, harnessing emotion and intent cues in visual data to enhance information systems, and about AI for video streaming: methods, systems, and governance for real-time ingestion, analysis, and decisioning over continuous video flows.
The goal is to bridge computer-vision analytics and affective models (extracting sentiment, engagement, trustworthiness, etc.) to design intelligent, human-centered systems that inform decision-making and adaptive services. The special issue is interested in methods that fuse visual, auditory, and textual data (multimodal analysis) to detect emotion, sentiment, engagement, gaze, gesture, posture, prosody and language in real time. Such techniques enable insights from live-streamed user interactions or "experience mining" in marketing and e-commerce. Deep-learning advances enable machines to recognize facial expressions, gestures, tone of voice and other cues to infer sentiment, engagement and intent. Recent work demonstrates that Convolutional Neural Networks combining video and audio input can predict viewers' emotions with high accuracy, achieving around a 75% AUC on real-world video-ads data. Other research has integrated language and vision: one study built an "Affective Mimicry Index" from CEO interview videos, using computer vision together with large language models, and found that higher video-derived empathy scores correlate with better firm performance. These examples show how multimodal video analytics can extract rich affective signals, from basic expressions to higher-level trust and empathy cues. The aim is to bring these advances into intelligent information systems that are explainable, fair, and aligned with human values.