OmniFall: From Staged Through Synthetic to Wild
A Unified Multi-Domain Dataset for Robust Fall Detection

David Schneider1,†, Zdravko Marinov1, Moritz Mistol1, Zeyun Zhong1, Alexander Jaus1, Rodi Düger1, Rafael Baur1, M. Saquib Sarfraz2, Rainer Stiefelhagen1
1Karlsruhe Institute of Technology 2Mercedes-Benz Tech Innovation
†Corresponding author: david.schneider∂kit.edu
ECCV 2026

Come and see us in Malmö. We present OmniFall in Poster Session 3, at 10:30 CEST, in ExHall at board #445. Session page

Poster frame of the OmniFall ECCV 2026 presentation

Five-minute ECCV 2026 presentation of OmniFall.

Abstract

Visual fall detection models are usually trained on comparatively small, staged datasets. Such data lacks diversity and evaluation protocols differ from paper to paper. Evaluation across datasets and within real world environments is difficult.

For this reason we present OmniFall, a unified benchmark of 15k videos (80 hours) with frame-level annotations in a unified 16-class taxonomy. It spans three domains: OF-Staged comprises eight staged datasets with cross-subject and cross-view splits; OF-Synthetic adds 12k synthetically generated videos (17 h) and OF-In-the-Wild provides a test-only set of genuine accidents.

We evaluate fine-tuned models as well as much larger multimodal LLMs in zero-shot settings. On in-the-wild fall events, both do comparably well. The critical continuous fallen state is where they part: zero-shot models keep confusing fallen with lying, whereas models fine-tuned on synthetic data with explicit fallen-state scenes do substantially better. We release the unified annotations, the synthetic data, and the in-the-wild test set to foster the development of fall and fallen-state detectors for uncontrolled environments.

Get the Data

We publish the annotations, the splits and the OF-Synthetic videos on the Hugging Face Hub. The videos of the other nine sub-datasets were recorded by other research groups, whose licences do not permit us to pass them on, so each has to be downloaded from the site that hosts it.

The omnifall python package does that for you. It downloads the annotations, retrieves each sub-dataset from its original source, converts the videos into one common layout and encoding and provides dataloaders which return decoded clips ready for training.

pip install omnifall

Every sub-dataset except CMDFall is fetched without automatically. CMDFall itself has to be requested by its authors, so you have to provide the local path to the data.

omnifall status            # list per-dataset preparation status
omnifall prepare of-syn    # download and convert one subset
omnifall verify le2i       # check a local copy

Data loading with PyTorch:

import omnifall
from torch.utils.data import DataLoader

parts  = omnifall.load_video_dataset("of-sta-to-all-cs", num_frames=16)
loader = DataLoader(parts["train"], batch_size=8,
                    collate_fn=omnifall.collate_fn)

batch = next(iter(loader))
batch["pixel_values"].shape   # (8, 16, 3, 224, 224)

Data loading and training with Huggingface datasets and transformers:

from transformers import Trainer, TrainingArguments

parts = omnifall.trainer_dataset("le2i-cs", model_name="MCG-NJU/videomae-base")
trainer = Trainer(
    model=omnifall.load_model("MCG-NJU/videomae-base"),
    train_dataset=parts["train"],
    eval_dataset=parts["validation"],
    data_collator=omnifall.collate_fn,
    compute_metrics=omnifall.compute_metrics,
    args=TrainingArguments(output_dir="out", remove_unused_columns=False),
)
trainer.train()

The tensor layout matches the input format of Hugging Face video models. Temporal sampling is randomised for training and deterministic for evaluation, which makes reported results reproducible.

Preparing CMDFall

CMDFall is the one component that cannot be downloaded automatically, because its authors grant access individually. It is also the largest staged component, at 7.12 h and 6,026 segments, so most users will want it. Four steps:

1. Request access. Follow the instructions on the CMDFall site : Write to Prof. Trần Thị Thanh Hải at the MICA institute via thanh-hai.tran@mica.edu.vn / hai.tranthithanh1@hust.edu.vn, stating your institution and your intended research use.

2. Download the RGB recordings. OmniFall uses only the RGB stream, distributed as the colors directory. The other modalities are not needed.

3. Unpack it. The archive expands to colors/S{s}P{p}K{k}.avi, where S is the subject, P the performance and K the Kinect index. This is already the layout OmniFall expects, so no renaming is required. Place the unpacked colors directory in the download directory that omnifall sources cmdfall prints, which defaults to ~/.cache/omnifall/downloads/cmdfall/unpacked.

4. Convert and verify. From here the procedure is identical to every other component:

omnifall prepare cmdfall                     # from the download directory
omnifall prepare cmdfall --archive <path>    # or from anywhere else
omnifall verify cmdfall                      # expects 384 of 384 videos

The sources are 20 fps MJPEG, which the package re-encodes to H.264.

Dataset Overview

OmniFall comprises three domains under a single taxonomy and a common set of evaluation protocols, so that a model can be trained on one domain and evaluated on another. Full videos are densely labelled, which supports both clip classification and timeline segmentation.

OF-Staged contains eight public staged datasets, manually re-annotated into a common label space, with cross-subject (CS) and cross-view (CV) splits for each. It provides approximately 14 h of single-view footage, which extends to 61.5 h when synchronised views are included.

OF-Synthetic contains 12k generated videos (17 h) covering demographics and environments that are not represented in the staged data.

OF-In-the-Wild provides a test-only set of genuine accidents (2.65 h) curated from OOPS. It is excluded from training and is used to measure generalisation to uncontrolled conditions.

Example videos from the OmniFall benchmark
Figure 1: Example videos from the OmniFall benchmark showcasing the diversity of camera views, subjects, and environments.

The 16-class taxonomy separates transient actions such as fall, sit down and stand up from the static states that follow them, such as fallen, sitting and standing.

The separation of fall from fallen is clinically motivated. A fall lasts about one second, whereas the fallen state can persist for hours, and these prolonged "long-lie" episodes are associated with the severe medical outcomes. A detector that responds only to the impact therefore misses the condition that actually requires intervention, and the experiments below show that distinguishing fallen from lying is where current models are weakest.

Example videos from the OmniFall benchmark
Figure 2: Cross-dataset compatible annotations for CMDFall, with action segment predictions from MS-TCN++ trained on OmniFall.
Example videos from the OmniFall benchmark
Figure 3: Segmented share of each label within datasets and total single view duration.

OF-Synthetic

Staged datasets are recorded with few volunteers in few rooms, which limits the diversity they can provide. OF-Synthetic addresses this with 12k five-second clips generated using Wan 2.2 at 1280×720 and 16 fps, each depicting a distinct person in a distinct scene.

Each clip is generated from a structured prompt whose axes vary independently. Demographics cover six age groups and the seven OMB SPD-15 race and ethnicity categories, together with body type and gender, across more than one thousand environments. Each fall combines one of 30 mechanisms with a direction and a recovery attempt, and camera placements range from eye level to elevated CCTV and top-down views.

This supplies wider demographic and environmental coverage than staged recordings and controlled generation of explicit fallen-state footage.

Generated data introduces its own biases. A generative model reflects the distribution of its training data, so faithful representation of every attribute cannot be guaranteed, particularly for underrepresented groups. Users of OF-Syn are asked to evaluate the resulting models for remaining biases which might be propagated in training.

What We Found

Four observations from our experiments. The paper reports the full protocol and results.

  • Detecting the fallen state is harder than detecting the fall itself, and it is the clinically decisive class.
  • Zero-shot multimodal LLMs such as Qwen3-VL match fine-tuned video models such as VideoMAE on in-the-wild falls, but not on the fallen state, which they confuse with lying.
  • Training on OF-Synthetic alone generalises to in-the-wild data better than training on OF-Staged, for both classification and timeline segmentation.
  • A substantial staged-to-wild gap remains for every model we evaluated.

Datasets

OmniFall unifies the following sub-datasets. The final column states whether omnifall prepare can obtain the videos without manual steps.

Dataset Duration Subjects Views/Rooms License omnifall prepare
CMDFall 7h 25m 50 7 synchronized views Permission from authors manual download
UP Fall 4h 35m 17 2 synchronized views Permission from authors yes
Le2i 47m 9 6 different rooms, 1 view each CC-BY-NC-SA 3.0 yes
GMDCSA24 21m 4 3 rooms, 1 view each MIT yes
CAUCAFall 16m 10 1 room, 1 view CC BY 4.0 yes
EDF 13m 5 2 views synchronized CC BY 4.0 yes
OCCU 14m 5 2 views not synchronized CC BY 4.0 yes
MCFD 12m 1 8 views synchronized Permission from authors yes
OF-In-the-Wild (OOPS) 2h 39m 818 videos uncontrolled, 1 view each CC BY-NC-SA 4.0 yes
OF-Synthetic 16h 53m 12,000 unique 1,000+ environments CC-BY-NC 4.0 yes

OF-Staged comprises approximately 14 hours of unique single-view recordings, which extends to 61.5 hours once synchronised concurrent views are counted. It covers 101 subjects across 29 camera views, and therefore requires generalisation across human appearance, environmental conditions and camera perspective.

BibTeX

Please do not only cite the OmniFall paper, but also the original papers of the sub-datasets you used. The following BibTeX entries are provided for your convenience.
@misc{omnifall,
      title={OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection}, 
      author={David Schneider and Zdravko Marinov and Moritz Mistol and Zeyun Zhong and Alexander Jaus and Rodi Düger and Rafael Baur and M. Saquib Sarfraz and Rainer Stiefelhagen},
      year={2025},
      eprint={2505.19889},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2505.19889}, 
},
@inproceedings{omnifall_cmdfall,
  title     = {A multi-modal multi-view dataset for human fall analysis and preliminary investigation on modality},
  author    = {Tran, Thanh-Hai and Le, Thi-Lan and Pham, Dinh-Tan and Hoang, Van-Nam and Khong, Van-Minh and Tran, Quoc-Toan and Nguyen, Thai-Son and Pham, Cuong},
  booktitle = {2018 24th International Conference on Pattern Recognition (ICPR)},
  pages     = {1947--1952},
  year      = {2018},
  doi       = {10.1109/ICPR.2018.8546308}
}

@article{omnifall_up-fall,
  title   = {{UP}-Fall Detection Dataset: A Multimodal Approach},
  author  = {Martínez-Villaseñor, Lourdes and Ponce, Hiram and Brieva, Jorge and Moya-Albor, Ernesto and Núñez-Martínez, José and Peñafort-Asturiano, Carlos},
  journal = {Sensors},
  volume  = {19},
  number  = {9},
  pages   = {1988},
  year    = {2019},
  doi     = {10.3390/s19091988}
}

@article{omnifall_le2i,
  title   = {Optimized spatio-temporal descriptors for real-time fall detection: comparison of support vector machine and Adaboost-based classification},
  author  = {Charfi, Imen and Miteran, Johel and Dubois, Julien and Atri, Mohamed and Tourki, Rached},
  journal = {Journal of Electronic Imaging},
  volume  = {22},
  number  = {4},
  pages   = {041106},
  year    = {2013},
  doi     = {10.1117/1.JEI.22.4.041106}
}

@article{omnifall_gmdcsa,
  title   = {{GMDCSA}-24: A dataset for human fall detection in videos},
  author  = {Alam, Ekram and Sufian, Abu and Dutta, Paramartha and Leo, Marco and Hameed, Ibrahim A.},
  journal = {Data in Brief},
  volume  = {57},
  pages   = {110892},
  year    = {2024},
  doi     = {10.1016/j.dib.2024.110892}
}

@article{omnifall_cauca,
  title   = {Dataset for human fall recognition in an uncontrolled environment},
  author  = {Guerrero, José Camilo Eraso and España, Elena Muñoz and Añasco, Mariela Muñoz and Lopera, Jesús Emilio Pinto},
  journal = {Data in Brief},
  volume  = {45},
  pages   = {108610},
  year    = {2022},
  doi     = {10.1016/j.dib.2022.108610}
}

@incollection{omnifall_edf_occu,
  title     = {Evaluating Depth-Based Computer Vision Methods for Fall Detection under Occlusions},
  author    = {Zhang, Zhong and Conly, Christopher and Athitsos, Vassilis},
  booktitle = {Advances in Visual Computing},
  volume    = {8888},
  pages     = {196--207},
  publisher = {Springer International Publishing},
  year      = {2014},
  doi       = {10.1007/978-3-319-14364-4_19}
}

@misc{omnifall_mcfd,
  title        = {Multiple cameras fall data set},
  author       = {Auvinet, Edouard and Rougier, Caroline and Meunier, Jean and St-Arnaud, Alain and Rousseau, Jacqueline},
  howpublished = {Technical Report 1350, DIRO, Université de Montréal},
  year         = {2010},
  url          = {https://web.archive.org/web/20250615085148/https://www.iro.umontreal.ca/~labimage/Dataset/technicalReport.pdf}
}

@inproceedings{omnifall_oops,
  title     = {Oops! Predicting Unintentional Action in Video},
  author    = {Epstein, Dave and Chen, Boyuan and Vondrick, Carl},
  booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  pages     = {916--926},
  year      = {2020},
  doi       = {10.1109/CVPR42600.2020.00100}
}

Acknowledgements

This work has been supported by the Carl Zeiss Foundation through the JuBot project as well as by funding from the pilot program Core-Informatics of the Helmholtz Association (HGF). The authors acknowledge support by the state of Baden-Württemberg through bwHPC. Experiments were performed on the HoreKa supercomputer funded by the Ministry of Science, Research and the Arts Baden-Württemberg and by the Federal Ministry of Education and Research.