Artificial Intelligence, Affective Computing and Video Analytics for Intelligent Information Systems

Editors

  • Andrea Generosi, Pegaso University
  • Luigi Gallo, Pegaso University
  • Valerio De Luca, Pegaso University
  • Maura Mengoni, Marche Polytechnic University
  • Josef Spjut, NVIDIA
  • Lucio De Paolis, University of Salento

Description

The past decade has seen explosive growth in video data (live and on-demand), driven by ubiquitous cameras, mobile devices, and online platforms. From city CCTV and body-cams to retail, sports, tele-health, and social live commerce, organizations now operate end-to-end video streaming pipelines at massive scale. In parallel, advances in deep learning, spanning video transformers, self-supervised pretraining, and multimodal large models, have transformed what intelligent systems can infer from continuous visual data, moving beyond static clips to long-horizon, low-latency understanding of events, activities, and interactions.

Affective computing adds a new dimension by interpreting emotional and social signals in video and multimedia. For example, research shows that portraying positive emotions and trust cues in video significantly increases viewer engagement. By combining facial expression, gesture, and acoustic analysis, advanced AI can infer user sentiment and intent. These capabilities enrich social-media intelligence, recommender systems, and customer journey analysis.

This special issue seeks cutting-edge work at the intersection of video analytics and affective computing, harnessing emotion and intent cues in visual data to enhance information systems, and about AI for video streaming: methods, systems, and governance for real-time ingestion, analysis, and decisioning over continuous video flows.

The goal is to bridge computer-vision analytics and affective models (extracting sentiment, engagement, trustworthiness, etc.) to design intelligent, human-centered systems that inform decision-making and adaptive services. The special issue is interested in methods that fuse visual, auditory, and textual data (multimodal analysis) to detect emotion, sentiment, engagement, gaze, gesture, posture, prosody and language in real time. Such techniques enable insights from live-streamed user interactions or "experience mining" in marketing and e-commerce. Deep-learning advances enable machines to recognize facial expressions, gestures, tone of voice and other cues to infer sentiment, engagement and intent. Recent work demonstrates that Convolutional Neural Networks combining video and audio input can predict viewers' emotions with high accuracy, achieving around a 75% AUC on real-world video-ads data. Other research has integrated language and vision: one study built an "Affective Mimicry Index" from CEO interview videos, using computer vision together with large language models, and found that higher video-derived empathy scores correlate with better firm performance. These examples show how multimodal video analytics can extract rich affective signals, from basic expressions to higher-level trust and empathy cues. The aim is to bring these advances into intelligent information systems that are explainable, fair, and aligned with human values.

Potential topics

  • Scalable video understanding: algorithms and architectures for analyzing large-scale video streams (e.g. in surveillance, social media or enterprise) to support decision-making and business analytics.
  • Emotion and sentiment analysis: detection of affective and engagement cues in video (e.g. customer-journey recordings, advertisement viewing, user-generated content).
  • Multimodal fusion: methods that combine vision with audio, text or other modalities to enrich social-media intelligence and contextual analysis.
  • Streaming-native AI and MLOps: online inference under tight latency budgets, stream processing/windowing, drift detection, A/B testing and safe rollouts, cost–latency–accuracy trade-offs, and end-to-end observability for production pipelines.
  • Video(-language) models for streaming: multimodal LLMs/VLMs, memory and token compression for long-horizon video, retrieval-augmented streaming
  • Real-time and distributed architectures: edge/cloud or federated systems for real-time video analytics at scale, including considerations of latency, bandwidth and privacy.
  • Explainability, fairness and compliance: approaches to make affective video AI transparent and trustworthy, addressing ethical, legal and regulatory challenges.
  • Adaptive learning: techniques for domain adaptation, continual learning or few-shot learning to handle evolving video streams and reduce the need for large labeled datasets.
  • Case studies and applications: empirical studies in domains such as healthcare (e.g. remote patient monitoring), finance (e.g. behavioral analytics), education (e.g. XR/serious games environments), and smart environments.