The SPHERE human skeleton movements dataset was created using a Kinect camera, that measures distances and provides a depth map of the scene instead of th…
action, behavior, depth, human, kinect, motion, movement, skeleton, videoChairGest is an open challenge / benchmark. The task consists in spotting and recognizing gestures from multiple synchronized sensors: 1 Kinect and 4 Xse…
benchmark, detection, gesture, human, kinect, recognitionThe Multicamera Human Action Video Data (MuHAVi) Manually Annotated Silhouette Data (MAS) are two datasets consisting of selected action sequences for th…
action, background, behavior, human, segmentation, videoThe PETS 2009 dataset contains 3 parts showing multi-view sequences containing pedestrians walking in an outdoor environment. The parts are used for perso…
detection, frontview, human, occlusion multitarget, outdoor, overlap, pedestrian, trackingThe High Definition Analytics (HDA) dataset is a multi-camera High-Resolution image sequence dataset for research on High-Definition surveillance: Pedestr…
benchmark, camera, detection, high-definition, human, indoor, lisbon, multiview, network, pedestrian, re-identification, surveillance, tracking, videoThe Shefeld Kinect Gesture (SKIG) dataset contains 2160 hand gesture sequences (1080 RGB sequences and 1080 depth sequences) collected from 6 subjects. Al…
action, depth, gesture, human, illumination, kinect, recognitionThe QMUL Junction dataset is a busy traffic scenario for research on activity analysis and behavior understanding. Video length: 1 hour (90000 frames)…
behavior, counting, crowd, detection, motion, pedestrian, tracking, videoGroup emotion recognition in images - Happiness Intensity labels for group of people in images. The images have been collected from Flickr using keyword s…
behavior, emotion, facial expression, flickr, group, human, wildThe UrbanStreet dataset used in the paper can be downloaded here [188M] . It contains 18 stereo sequences of pedestrians taken from a stereo rig mounted o…
detection, human, multitarget, pedestrian, recognition, segmentation, tracking, urban, videoThe TUG (Timed Up and Go test) dataset consists of actions performed three times by 20 volunteers. The people involved in the test are aged between 22 and…
accelerometer, action, depth image processing - tug, human, kinect, recognition, time, video, wearableMICCAI 2015 Challenge on Liver Ultrasound Tracking Munich, October 9, 2015 (Full Day) Outline Ultrasound (US) imaging is a widely used medical imaging…
benchmark, human, liver, medical, organ, real, therapy, tracking, ultrasoundWelcome to the homepage of the gvvperfcapeva datasets. This site serves as a hub to access a wide range of datasets that have been created for projects of…
action, depth, face, human, mesh, multiview, pose, reconstruction, tracking, videoA 66 stereo pairs dataset with their subpixel ground truths. The construction and improvement of algorithms for subpixel stereovision requires very prec…
3d, depth, groundtruth, noise, pointcloud, stereo, stereovision, subpixelIt is composed of ADL (activity daily living) and fall actions simulated by 11 volunteers. The people involved in the test are aged between 22 and 39, wit…
accelerometer, action, depth, fall detection - adl, human, kinect, recognition, video, wearableThe dataset consist of the about 50 hours obtained from kindergarten surveillance videos. Dataset, totally approximately 100 videos sequences (1000GB, 50 …
action, background, behavior, human, segmentation, video surveillanceThe Video Summarization (SumMe) dataset consists of 25 videos, each annotated with at least 15 human summaries (390 in total). The data consists of videos…
action, benchmark, event, groundtruth, human, summary, videoA large dataset of geotagged face images collected from Flickr. The zip file contains text files containing urls of the images. Face2GPS: Estimating Geo…
age, classification, face, gender, geotagged, human, localizationThe Video Segmentation Benchmark (VSB100) provides ground truth annotations for the Berkeley Video Dataset, which consists of 100 HD quality videos divide…
benchmark, groundtruth, motion, object, pedestrian, segmentation, tracking, videoShakeFive2 A collection of 8 dyadic human interactions with accompanying skeleton metadata. The metadata is frame based xml data containing the skeleton…
human, interaction, kinect, videoThe MSR RGB-D Dataset 7-Scenes dataset is a collection of tracked RGB-D camera frames. The dataset may be used for evaluation of methods for different app…
depth, kinect, location, reconstruction, tracking, videoThe Microsoft Research Cambridge-12 Kinect gesture dataset consists of sequences of human movements, represented as body-part locations, and the associate…
action, gesture, human, kinect, recognitionThe CHALEARN Multi-modal Gesture Challenge is a dataset +700 sequences for gesture recognition using images, kinect depth, segmentation and skeleton data.…
action, depth, gesture, human, illumination, kinect, recognition, segmentation, skeletonThe Freiburg-Berkeley Motion Segmentation Dataset (FBMS-59) is an extension of the BMS dataset with 33 additional video sequences. A total of 720 frames i…
benchmark, groundtruth, motion, object, pedestrian, segmentation, tracking, videoSome datasets and evaluation tools are provided on this page for four different computer vision and computer graphics problems. Population counting Lin…
3d, counting, crowd, detection, groundtruth, line, network, object, pedestrian, pointcloud, reconstruction, road, surface, urbanThe Inria Aerial Image Labeling addresses a core topic in remote sensing: the automatic pixelwise labeling of aerial imagery (link to paper). Dataset fe…
aerial, building, city, footprint, groundtruth, house, segmentation, semantic, urbanMIT traffic data set is for research on activity analysis and crowded scenes. It includes a traffic video sequence of 90 minutes long. It is recorded by a…
trackingThe INRIA People dataset from Navneet Dalal and Bill Triggs [DalalCVPR2005] consists of training and testing data. The training contains 1805 images and X…
boundingbox, frontview, human, object detection, pedestrian, sideviewThe Buffy dataset contains images selected from the TV series, Buffy: the Vampire Slayer. We select a set of 452 images from the first two episodes for tr…
buffy, human, movie, object detection, segmentationThe Quad 6K dataset is a Structure-from-Motion dataset taken at Arts Quad at Cornell University campus and consists of 6514 images with ground truth posit…
3d gps, 3d reconstruction, groundtruth, landmark, sfm, urbanRobust Multi-Person Tracking from Mobile Platforms In all cases, data was recorded using a pair of AVT Marlins F033C mounted on a chariot respectively a…
color, pedestrian, sequence, trackingThe Ford Car dataset is joint effort of Pandey et al. (for collecting images, Lidar points, calibration etc.) and us (for annotation of 2D and 3D objects)…
3d, car, detection, groundtruth, lidar, sfmThe domain-specific personal videos highlight dataset from the paper [1] describes a fully automatic method to train domain-specific highlight ranker for…
action, domain, human, recognition, saliency, summarization, video, wearableDataset contains 1000 images of 100 persons, with 10 images per person and is freely available. All images were acquired by cropping ears from images from…
biometry, ear, human, lighting, pedestrian, person, recognition3 datasets: PTZ Tracking, Thermal-visible registration, Single object tracking
pedestrian, ptz, thermal, trackingSince the publicly available face image datasets are often of small to medium size, rarely exceeding tens of thousands of images, and often without age in…
age, biometry, detection, face, imdb, recognition, wikipediaThe PIROPO database (People in Indoor ROoms with Perspective and Omnidirectional cameras) comprises multiple sequences recorded in two different indoor ro…
detection, fisheye, human, indoor, omnidirectional, people, perspective, room, surveillanceThe database contains, for each of the 100 examples: (1) the uncompressed frames, up to the 10th frame after the appearance of the 8th cell; (2) a text fi…
biology, cell, circle, mouse, tracking, trajectoryDataset A (former NLPR Gait Database) was created on Dec. 10, 2001, including 20 persons. Each person has 12 image sequences, 4 sequences for each of the …
action, biometry, classification, foot, gait, human, motion, pressure, recognitionThe NYU-Depth data set is comprised of video sequences from a variety of indoor scenes as recorded by both the RGB and Depth cameras from the Microsoft Ki…
depth, kinect, label, reconstruction, semantic segmentationThis dataset consist 51 oral presentation recorded with 2 ambient visual sensor (web-cam), 3 First Person View (FPV) cameras (1 on presenter and 2 on rand…
analysis, kinect, multi-sensor, presentation, quality, videoThe UCF Person and Car VideoSeg dataset consists of six videos with groundtruth for video object segmentation. Surfing, jumping, skiing, sliding, big ca…
camera, groundtruth, model, motion, object, segmentation, videoThe Robotic 3D Scan Repository from Osnabrueck contains 23 different datasets showing a veriaty of 3D scans for objects, humans, cities, university campus…
3d, aerial, bremen, city, germany, heat, human, laser, lidar, osnabrueck, reconstruction, scan, urbanFor the first few decades of the fields existence, computer vision has been focused on algorithmic, logical approaches to perception. But it was only with…
3d, depth, indoor, kinect, object, recognition, reconstructionThe mirror symmetry database contains 176 single-symmetry and 63 multyple-symmetry images (.png files) with accompanying ground-truth annotations (.mat fi…
detection, groundtruth, mirror, symmetryThe Pittsburgh Fast-food Image dataset (PFID) consists of 4545 still images, 606 stereo pairs, 3033600 videos for structure from motion, and 27 privacy-pr…
classification, food, laboratory, real, recognition, reconstruction, videoThis ETHZ CVL RueMonge 2014 dataset used for 3D reconstruction and semantic mesh labelling for urban scene understanding. It was first published in [1] …
3d, architecture, benchmark, classification, code, mesh, outdoor, paris, pointcloud, recognition, reconstruction, segmentation, semantic, source, urbanJPL First-Person Interaction dataset (JPL-Interaction dataset) is composed of human activity videos taken from a first-person viewpoint. The dataset parti…
action, human, interactive, motion, recognition, videoThe Our Database of Faces (ORL) dataset contains ten different images of each of 40 distinct subjects. For some subjects, the images were taken at differe…
expression, face, human, illumination, recognitionThe Annotated Facial Landmarks in the Wild (AFLW) consists of a large-scale collection of annotated face images gathered from the web, exhibiting a large …
age, annotation, detection, face, landmark, poseThe Multi-FoV synthetic datasets are two synthetic scenes (vehicle moving in a city, and flying robot hovering in a confined room). For each scene, three …
blender, camera, fov, groundtruth, odometry, synthetic, visualThis data set comprises 144 images of an edge profile cutting head of a milling machine. The head tool contains a total of 30 cutting inserts. The cutting…
cutting, edge, head, inserts, localization, milling, monitoring, object, profile, tool, tools, wearMultispectral Imaging (MSI) datasets were acquired using IRIS II which is a lightweight portable system comprising of a high resolution camera, a novel fi…
alignment, groundtruth, illumination, matching, multi-spectral, registration, wavelengthThe PETS 2016 IPATCH dataset contains a set of fourteen multi camera recordings (visible, themal) collected off the coast of Brest, France, in collaborati…
boat, detection, gps, maritime, multimodal, radar, thermal, tracking, vessel, visibleThe tracking environment consists of multiple 3D range sensors, covering an area of about 900 m2, in the "ATC" shopping center in Osaka, Japan.
trackingThe BEOID dataset includes object interactions ranging from preparing a coffee to operating a weight lifting machine and opening a door. The dataset is re…
3d, egocentric, interaction, object, pose, tracking, videoHallway Corridor - Multiple Camera Tracking: An indoor camera network dataset with 6 cameras (contains ground plane homography).
trackingAn Annotated Dataset For Near-Duplicate Detection In Personal Photo Collections Managing photo collections involves a variety of image quality assessmen…
copyright, detection, duplicate, groundtruth, retrievalThe Stanford 40 Actions dataset contains images of humans performing 40 actions. In each image, we provide a bounding box of the person who is performing …
action, boundingbox, detection, human, recognitionThe TU Berlin Multi-Object and Multi-Camera Tracking Dataset (MOCAT) is a synthetic dataset to train and test tracking and detection systems in a virtual …
animal, detection, evaluation, multi-class, multi-view, pedestrian, synthetic, tracking, vehicleThe SegTrack dataset consists of six videos (five are used) with ground truth pixelwise segmentation (6th penguin is not usable). The dataset is used for …
camera, flow, groundtruth, model, motion, object, optical, proposal, segmentation, stationary, videoThe set was recorded in Zurich, using a pair of cameras mounted on a mobile platform. It contains 12'298 annotated pedestrians in roughly 2'000 frames.
trackingThe VOT2016 pixel-wise annotations dataset contains pixel-wise per-frame annotations for sequences from VOT2016 dataset. The annotation is in a form of BW…
annotation, mask, object, segmentation, tracking, visualThe Traffic Video dataset consists of X video of an overhead camera showing a street crossing with multiple traffic scenarios. The dataset can be downlo…
detection, overhead, road, tracking, traffic, urban, video, viewThe Malaya Abrupt Motion (MAMo) dataset is targeted for visual tracking, particularly for abrupt motion tracking. It was collected from publicly accessibl…
abrupt motion tracking, tracking, visual trackingThe Mall dataset was collected from a publicly accessible webcam for crowd counting and profiling research. Ground truth: Over 60,000 pedestrians were …
counting, crowd, detection, indoor, pedestrian, tracking, video, webcamThe PASCAL VOC is augmented with segmentation annotation for semantic parts of objects. For example, for the person category, we provide segmentation mask…
detection, human, object, part, pascal, pedestrian, recognition, segmentation, semanticThe ICG Lab 6 (Multi-Camera Multi-Object Tracking) dataset contains 6 indoor people tracking scenarios recorded at our laboratory using 4 static Axis P134…
calibration, camera, detection, evaluation, graz, laboratory, multiview, object, pedestrian, segmentation, trackingWe share our omnidirectional and panoramic image dataset (with annotations) to be used for human and car detection. Please reach through: http://cvrg.iyt…
car, detection, human, omnidirection, panorama, recognitionThe crowd datasets are collected from a variety of sources, such as UCF and data-driven crowd datasets. The sequences are diverse, representing dense crow…
anomaly, crowd, detection, human, pedestrian, scene, understanding, videoThe NYU-Depth V2 data set is comprised of video sequences from a variety of indoor scenes as recorded by both the RGB and Depth cameras from the Microsoft…
depth, kinect, label, reconstruction, semantic segmentationLow-resolution RGB videos + ground truth trajectories from multiple fixed and moving cameras monitoring the same scenes (indoor and outdoor) to improve ob…
trackingThe dataset captures 25 people preparing 2 mixed salads each and contains over 4h of annotated accelerometer and RGB-D video data. Annotated activities co…
action, activity, classification, detection, recognition, tracking, videoCollected in a clothing store. Captured with Kinect (640*480, about 30fps)
detection, trackingWe introduce the Shelf dataset for multiple human pose estimation from multiple views. In addition we annotate the body joints in the Campus dataset from …
3d, capture, estimation, human, motion, multiple, pose, viewThe CVC Partial Occlusion Virtual Pedestrian datasets (CVC-01 to CVC-06) cover a range of scenarios of occluded pedestrians generated in a virtual and rea…
classification, detection, occlusion, pedestrian, synthetic, tracking, urbanThe 1DSfM Landmarks is a collection of community-based image reconstruction by Kyle Wilson and is comprised of 14 datasets with comparison to bundler grou…
3d, benchmark, city, groundtruth, landmark, reconstruction, urbanCOCO-Stuff augments the COCO dataset with pixel-level stuff annotations for 10,000 images. These annotations can be used for scene understanding tasks lik…
annotation, benchmark, captioning, coco, groundtruth, segmentation, semantic, stuff, thingsThis dataset consists of more than 22,000 images of 24 people which are captured by 16 cameras installed in a shopping mall "Shinpuh-kan". All images are …
trackingThe TUD Crossing dataset from Micha Andriluka, Stefan Roth and Bernt Schiele consists of 201 images with 1008 highly overlapping pedestrians with signific…
detection, multitarget, overlap, pedestrian, segmentation, sideview, tracking, urbanThe Landmark 1000 or 1k dataset is a collection of the top 1000 popular flickr landmarks mined from flickr. It is maintained by Noah Snavely and publish…
3d, estimation, landmark, location, pointcloud, pose, reconstruction, worldThe KTH Multiview Football dataset contains 771 images of football players includes images taken from 3 views at 257 time instances 14 annotated body join…
camera, detection, game, multitarget, multiview, object, outdoor, pedestrian, pose, recognition, soccer, trackingThe Salient Montages is a human-centric video summarization dataset from the paper [1]. In [1], we present a novel method to generate salient montages f…
human, montage, saliency, summarization, video, wearableLASIESTA is composed by many real indoor and outdoor sequences organized in different categories, each of one covering a specific challenge in moving obje…
background, camera, challenge, dataset, detection, foreground, groundtruth, motion, object, stationary, subtractionWe collected a video dataset, termed ChokePoint, designed for experiments in person identification/verification under real-world surveillance conditions u…
clustering, detection, face, human, identification, multiview, pedestrian, real, recognition, sequence, surveillance, worldThe multi-modal/multi-view datasets are created in a cooperation between University of Surrey and Double Negative within the EU FP7 IMPART project. The …
3d, action, color, dynamic, emotion, face, human, indoor, lidar, model, multi-mode, multi-view, outdoor, rgbd, videoThe dataset consists of eight unique scenes in crowded spaces such as a university campus or the sidewalks of a busy street.
trackingAWS hosts a variety of public datasets that anyone can access for free. Previously, large datasets such as satellite imagery or genomic data have require…
amazon, biology, classification, deep, human, image, learning, recognition, resolution, satellite, segmentation, spaceParis-rue-Madame dataset contains 3D Mobile Laser Scanning (MLS) data from rue Madame, a street in the 6th Parisian district (France). The test zone conta…
3d, classification, laser, pointcloud, segmentation, semanticThe test sequences provide interested researchers a real-world multi-view test data set captured in the blue-c portals. The data is meant to be used for t…
action, camera, multiview, segmentation, trackingThis dataset package contains the software and data used for Detection-based Object Labeling on the RGB-D Scenes Dataset as implemented in the paper: De…
3d, depth, indoor, kinect, object, recognition, reconstructionThe Notre Dame de Paris dataset used for 3D SfM reconstruction and contains 715 images provided by Noah Snavely. There are also version for NotreDame b…
3d, 3d reconstruction, flickr, frontview, landmark, limited, paris, pointcloud, sfmThe Raw Food Texture database (RawFooT) has been specially designed to investigate the robustness of descriptors and classification methods with respect t…
food, textureDatabase contains 798 images of 114 persons, with 7 images per person and is freely available for research purposes. All images were taken in supervised c…
biometry, face, human, illumination, lighting, pedestrian, person, recognitionThe ICG Multi-Camera and Virtual PTZ dataset contains the video streams and calibrations of several static Axis P1347 cameras and one panoramic video from…
calibration, camera, crowd, detection, graz, multitarget, multiview, network, object, outdoor, panorama, pedestrian, tracking, videoThe ICG Multi-Camera datasets consist of Easy Data Set (just one person) Medium Data Set (3-5 persons, used for the experiments) Hard Data Set (crowd…
calibration, camera, detection, graz, indoor, multitarget, multiview, object, pedestrian, tracking, videoThe MOT Challenge is a framework for the fair evaluation of multiple people tracking algorithms. In this framework we provide: - A large collection of d…
3d, benchmark, benhttp://motchallenge.net/chmark, dataset, evaluation, multiple, pedestrian, people, surveillance, target, tracking, videoThe FaceScrub dataset comprises a total of 107818 unconstrained face images of 530 celebrities crawled from the Internet, with about 200 images per person…
celebrity, detection, face, human, people, recognitionWe present the 2017 DAVIS Challenge, a public competition specifically designed for the task of video object segmentation. Following the footsteps of othe…
benchmark, code, hd, object, quality, resolution, segmentation, tracking, video segmentation