Deep Learning for Videos: A 2018 Guide to Action Recognition
Summary of major landmark action recognition research papers till 2018
A curated list of action recognition and related area resources
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
Summary of major landmark action recognition research papers till 2018
Brief human action recognition literature survey of work published between 2014 and 2019.
J. Choi et al., NeurIPS2019. [project web] [code] [arXiv]
C. Feichtenhofer et al., ICCV2019. [code]
D. Ghadiyaram et al., arXiv2019.
D. Tran et al., arXiv2019.
R. Girdhar et al., arXiv2019.
B. Korbar et al., arXiv2019.
R. Girdhar et al., CVPR2019. [project web]
X. Wang et al., CVPR2019. [code] [project web]
AJ. Piergiovanni and M. S. Ryoo et al., CVPR2019.
C. Li et al., CVPR2019.
X. Liu et al., CVPR2019.
N. Hussein et al., CVPR2019.
J.-B. Alayrac et al., CVPR2019.
C.-Y. Wu. et al., CVPR2019. [code]
B. Zhou et al., ECCV2018. [code] [project web]
Codes for popular action recognition models, written based on pytorch, verified on the something-something dataset.
X. Wang and A. Gupta, ECCV2018.
K. Hara et al., CVPR2019. [code]
D. Tran et al., CVPR2018. [code] [PyTorch]
CY. Ma et al., CVPR 2018.
X. Wang et al., CVPR2018. [code]
S. Xie et al., arXiv2017.
D. Tran et al, arXiv2017. Note: Aka Res3D. [code]: In the repository, C3D-v1.1 is the Res3D implementation.
Z. Qui et al, ICCV2017. [code]
J. Carreira et al, CVPR2017. [code][PyTorch code], [another PyTorch code]
D. Tran et al, ICCV2015. [the official Caffe code] [project web] Note: Aka C3D. [Python Wrapper] Note that the official caffe does not support python wrapper. [TensorFlow], [TensorFlow + Keras], [Another TensorFlow Implemetation], [Keras C3D Project web]: [Keras code], [Pretrained weights].
A. Diba et al, CVPR2017.
C. Lea et al, CVPR 2017. [code]
G. Varol et al, TPAMI2017. [project web] [code]
L. Wang et al, arXiv 2016. [code]
C. Feichtenhofer et al, CVPR2016. [code]
K. Simonyan and A. Zisserman, NIPS2014.
M. Xu et al, ICCV2019. [code]
M. Xu et al, Neurips2021. [code]
3D ResNets for Action Recognition (CVPR 2018)
An open source toolkit for video understanding by OpenMMLab. It supports state-of-the-art models for action recognition, temporal action detection, and spatio-temporal detection in videos.
Efficient video reader for python
P. Pandey et al, AAAI 2020. [code]
M. Guo et al., ECCV2018.
A. Diba et al., CVPRW2018.
A. Diba et al., arXiv2017.
R. Girdhar and D. Ramanan, NIPS2017. [code]
Byeon et al, arXiv2017.
Y. Zhu et al, arXiv2017. [code]
H. Bilen et al, CVPR2016. [code] [project web]
J. Donahue et al, CVPR2015. [code] [project web]
L. Yao et al, ICCV2015. [code] note: from the same group of RCN paper “Delving Deeper into Convolutional Networks for Learning Video Representations"
L. Wang et al, BMVC2016.
B. Zhang et al, CVPR2016. [code]
L. Wang et al, CVPR2015. [code]
M. Li et al., CVPR2019.
C. Si et al., CVPR2019.
P. Zhang et al., TPAMI2019.
S. Yan et al., AAAI2018. [code]
Y. Tang et al., CVPR2018.
K. Thakkar et al., BMVC2018.
Yu-Wei Chao et al., CVPR2018
Phuc Nguyen et al., CVPR 2018
P. Lei and S. Todrovic., CVPR2018.
Shayamal Buch et al., BMVC 2017 [code]
Jiyang Gao et al., BMVC 2017 [code]
Kaufman et al., ICCV2017. [code]
Y. Zhao et al., ICCV2017. [code] [project web]
X. Dai et al., ICCV2017.
F. Heidarivincheh et al., arXiv2017.
Z. Shou et al, CVPR2017. [code]
S. Buch et al, CVPR2017. [code]
H. Xu et al, arXiv2017. [code] [project web] [PyTorch]
V. Escorcia et al, ECCV2016. [code] [raw data]
Y. Li et al, ECCV2016. Noe: RGB-D Action Detection
Z. Shou et al, CVPR2016. [code] Note: Aka S-CNN.
F. Heilbron et al, CVPR2016. [code] Note: Depends on C3D, aka SparseProp.
L. Wang et al, CVPR2016. [code] Note: The code is not a complete verision. It only contains a demo, not training. [project web]
S. Ma et al, CVPR2016.
S. Yeung et al, CVPR2016. [code] [project web] Note: This method uses reinforcement learning
G. Yu and J. Yuan, CVPR2015. Note: code for FAP is NOT available online. Note: Aka FAP.
P. Mettes et al, ICMR2015.
K. Soomro et al, ICCV2015.
R. Girdhar et al., ActivityNet Workshop, CVPR2018.
A. El-Nouby and G. Taylor, arXiv2018.
P. Weinzaepfel et al., arXiv2017.
K. Soomro and M. Shah, ICCV2017.
P. Mettes and C. G. M. Snoek, ICCV2017.
V. Kalogeiton et al, ICCV2017. [code] [project web]
R. Hou et al, ICCV2017. [project web]
M. Zolfaghari et al, ICCV2017. [project web]
H. Zhu et al., ICCV2017.
G. Singh et al, ICCV2017. [code]
S. Saha et al, ICCV2017.
F. Becattini et al, BMVC2017.
J. He et al, arXiv2017.
H. S. Behl et al, arXiv2017.
X. Peng and C. Schmid. ECCV2016. [code]
P. Mettes et al, ECCV2016.
S. Saha et al, BMVC2016. [code] [project web]
P. Weinzaepfel et al. ICCV2015.
W. Chen and J. Corso, ICCV2015.
G. Gkioxari and J. Malik CVPR2015. [code] [project web]
J. Gemert et al, BMVC2015. [code]
D. Oneata et al, ECCV2014. [code] [project web]
M. Jain et al, CVPR2014.
Y. Tian et al, CVPR2013. [code]
K. Soomro et al, ICCV2015.
G. Yu and J. Yuan, CVPR2015. Note: code for FAP is NOT available online. Note: Aka FAP.
G. Sigurdsson et al., CVPR2018. [code]
P. Parma and B. T. Morris. CVPR2019.
S. Manen et al., ICCV2017.
A. Canziani and E. Culurciello - arXiv2017. [code] [project web]
J. Shao et al, CVPR2016. [code]
, paper
, paper, [INRIA web] for missing videos
, paper, download toolkit
A dataset of unintentional action, paper
a large-scale dataset for comprehensive instructional video analysis, paper
, technical report
, technical report
Daily Action Localization in Youtube videos. Note: Weakly supervised action detection dataset. Annotations consist of start and end time of each action, one bounding box per each action per video.
, 20BN-SOMETHING-SOMETHING
Note: They provide a download script and evaluation code here .
, paper - First person and third person video aligned dataset
, paper - First person videos recorded in kitchens. Note they provide download scripts and a python library here
Large scale action recognition dataset.
Note: It overlaps with UCF-101 dataset.
Note: It overlaps with UCF-101 dataset.
Spatio-Temporal annotations
, annotation provided by THUMOS-14, and corrupted annotation list, UCF-101 corrected annotations and different version annotaions. And there are also some pre-computed spatiotemporal action detection results
, note: the train/test split link in the official website is broken. Instead, you can download it from here.
Note: It overlaps with KTH datset.
Info and sample codes for "NTU RGB+D Action Recognition Dataset"
C. Vondrick et. al, IJCV2013. [code]
D. Mihalcik and D. Doermann, Technical report.
. Modern app to annotate objects in videos and images. It facilitates the development of an end-to-end machine learning pipeline encompassing the annotation/export/import of assets. Moreover, it could run as a native app or via web.
. Simple and standalone manual annotation web-app for image, audio and video. It runs in the web browser and does not require any installation or setup.
J. Dai et al., ICCV2017. [official code]
Open Source Object Detection Framework from Facebook AI Research. Includes Mask R-CNN, FPN, and etc. Caffe2 implementation.
K. He et al, [Detectron], [TensorFlow + Keras], [MXNet], [TensorFlow], [PyTorch] - State-of-the-art object detection/instance segmentation algorithm.
S. Ren et al, NIPS2015. [official MatCaffe code], [PyCaffe], [TensorFlow], [Another TF implementation] [Keras] - State-of-the-art object detector.
J. Redmon et al, CVPR2016. [official code], [TensorFLow] - Fast object detector.
J. Redmon and A. Farhadi, CVPR2017. [official code] - State-of-the-art object detector which can detect 9000 objects in realtime.
W. Liu et al, ECCV2016. [official PyCaffe code], [TensorFlow], [Keras] - State-of-the-art object detector with realtime processing speed.
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He and Piotr Dollár, Facebook AI Research FAIR & ICCV 2017.[Keras] - State-of-the-art object detector with realtime processing speed.
PyTorch based realtime and accurate pose estimation and tracking tool from SJTU.
R. Girdhar et al., arXiv2017.
Caffe based realtime pose estimation library from CMU.
Z. Cao et al, CVPR2017. [code] depends on the [caffe RT pose] - Earlier version of OpenPose from CMU.
[code] - Dense pose human estimation in the wild implemented in the Detectron framework.
M. Kocabas et al, ECCV2018. [code]
A. Mathis et al, Nature Neuroscience 2018. [code]
Activity detection in security camera videos. Runs through 2021. Hosted by NIST.
kdeldycke/awesome-falsehood
😱 Falsehoods Programmers Believe in
laoma2053/awesome-zhuiju-free
免费无广告的追剧资源指南,人工精选资源、每天检测资源有效性。收录在线影视、影视APP、网盘搜索、磁力BT、字幕、TVBox / 影视仓空壳软件/配置地址、IPTV直播源、会员拼团、影视相关开源项目。开源,社区共同维护。
ebu/awesome-broadcasting
A curated list of amazingly awesome open source resources related to broadcast technologies
44bits/awesome-opensource-documents
:blue_book: A curated list of awesome open source or open source licensed documents, guides, books.
chaofengc/Awesome-Image-Quality-Assessment
A comprehensive collection of IQA papers
PicoTrex/Awesome-Nano-Banana-images
A curated collection of fun and creative examples generated with Nano Banana & Nano Banana Pro🍌, Gemini-2.5-flash-image based model. We also release…