Journal of Computer Applications
Next Articles
Received:
Revised:
Accepted:
Online:
Published:
潘理虎1*,周飞1,樊光瑞1,张林梁2,李冰依1
通讯作者:
基金资助:
Abstract: Existing unsupervised Video Anomaly Detection (VAD) methods suffer from two critical limitations. First, high-level semantic information from human poses has not been effectively exploited, resulting in insufficient sensitivity to anomalies in behavioral logic. Second, memory modules are readily disturbed during training by multimodal feature interference and latent noise contamination in the data, which compromises the purity and robustness of prototype learning. To address these issues, a two-stage prototype memory framework based on Pose Awareness and cross-modal Motion Consistency Network (PAMC-Net) was proposed. A four-branch video decomposition scheme was constructed to separately model motion, scene, object, and pose information, and an independent prototype learning-and-freezing training paradigm was designed to preserve the purity of motion and pose prototypes while preventing mutual feature interference. In addition, a cross-modal consistency loss between pose and motion was introduced to incorporate physical plausibility constraints into unsupervised learning, enabling the detection of anomalous behaviors that violate human kinematics. Extensive experiments were conducted on four widely used benchmark datasets — CUHK Avenue, UCSD Ped1, UCSD Ped2, and ShanghaiTech — and the proposed framework achieves Area Under Curve (AUC) scores of 90.1%, 86.4%, 98.7%, and 76.6%, respectively. Compared with MS2NL (Multi-scale Spatiotemporal Normality Learning), an improvement of 1.7 percentage points in AUC is obtained by the proposed framework on the ShanghaiTech dataset. These results indicate that the proposed framework addresses the challenges posed by anomalous human behaviors in VAD effectively.
Key words: deep learning, Video Anomaly Detection (VAD), unsupervised learning, posture detection, cross-modal
摘要: 现有无监督视频异常检测(VAD)方法存在两方面的关键问题:一是未能有效利用人体姿态这一高层语义信息,导致对行为逻辑异常的敏感度不足;二是记忆模块在训练中易受多模态特征干扰及数据中隐含噪声污染,影响了原型学习的纯净性与鲁棒性。为解决上述问题,提出一种基于姿态感知和跨模态运动一致性的两阶段原型记忆框架(PAMC-Net)。所提框架构建运动、场景、对象与姿态的四路视频分解,并设计独立原型学习与冻结的训练范式,以确保运动与姿态原型的纯净性,避免特征互扰。此外,提出姿态、运动跨模态一致性损失,为无监督学习构建物理合理性约束,使模型能够检测出违背人体运动学的异常行为。在CUHK Avenue、UCSD Ped1、UCSD Ped2和ShanghaiTech这4个目前主流的基准数据集上,所提框架的曲线下面积(AUC)值分别达到了90.1%、86.4%、98.7%和76.6%;与MS2NL(Multi-Scale Spatiotemporal Normality Learning)方法相比,所提框架在ShanghaiTech数据集上的AUC值提升了1.7个百分点。可见,所提出的框架能够有效应对VAD中人体行为异常带来的挑战。
关键词: 深度学习, 视频异常检测, 无监督学习, 姿态感知, 跨模态
CLC Number:
TP391.4
潘理虎 周飞 樊光瑞 张林梁 李冰依. 面向姿态感知与运动一致性的视频异常检测[J]. 《计算机应用》唯一官方网站, DOI: 10.11772/j.issn.1001-9081.2026040460.
/ Recommend
Add to citation manager EndNote|Ris|BibTeX
URL: https://www.joca.cn/EN/10.11772/j.issn.1001-9081.2026040460