The SegTrack dataset consists of six videos (five are used) with ground truth pixelwise segmentation (6th penguin is not usable). The dataset is used for …
camera, flow, groundtruth, model, motion, object, optical, proposal, segmentation, stationary, videoThe UCF Person and Car VideoSeg dataset consists of six videos with groundtruth for video object segmentation. Surfing, jumping, skiing, sliding, big ca…
camera, groundtruth, model, motion, object, segmentation, videoThe KTH Multiview Football dataset contains 771 images of football players includes images taken from 3 views at 257 time instances 14 annotated body join…
camera, detection, game, multitarget, multiview, object, outdoor, pedestrian, pose, recognition, soccer, trackingScene Background Initialization (SBI) dataset The SBI dataset has been assembled in order to evaluate and compare the results of background initializati…
background, benchmark, change, detection, foreground, initializationThe Freiburg-Berkeley Motion Segmentation Dataset (FBMS-59) is an extension of the BMS dataset with 33 additional video sequences. A total of 720 frames i…
benchmark, groundtruth, motion, object, pedestrian, segmentation, tracking, videoThe ICG Multi-Camera datasets consist of Easy Data Set (just one person) Medium Data Set (3-5 persons, used for the experiments) Hard Data Set (crowd…
calibration, camera, detection, graz, indoor, multitarget, multiview, object, pedestrian, tracking, videoThe Video Segmentation Benchmark (VSB100) provides ground truth annotations for the Berkeley Video Dataset, which consists of 100 HD quality videos divide…
benchmark, groundtruth, motion, object, pedestrian, segmentation, tracking, videoThe ICG Multi-Camera and Virtual PTZ dataset contains the video streams and calibrations of several static Axis P1347 cameras and one panoramic video from…
calibration, camera, crowd, detection, graz, multitarget, multiview, network, object, outdoor, panorama, pedestrian, tracking, videoSome datasets and evaluation tools are provided on this page for four different computer vision and computer graphics problems. Population counting Lin…
3d, counting, crowd, detection, groundtruth, line, network, object, pedestrian, pointcloud, reconstruction, road, surface, urbanThe ICG Lab 6 (Multi-Camera Multi-Object Tracking) dataset contains 6 indoor people tracking scenarios recorded at our laboratory using 4 static Axis P134…
calibration, camera, detection, evaluation, graz, laboratory, multiview, object, pedestrian, segmentation, trackingPenn-Fudan Pedestrian Detection and Segmentation
background, detection, motion, pedestrian, segmentationThe GaTech VideoSeg dataset consists of two (waterski and yunakim?) video sequences for object segmentation. There exists no groundtruth segmentation an…
camera, model, motion, object, segmentation, videoBackground Models Challenge (BMC) is a complete dataset and competition for the comparison of background subtraction algorithms. The main topics concern:…
background, change, detection, modeling, motion, segmentation, surveillance, videoThis web page contains video data and ground truth for 16 dances with two different dance patterns. The style of dancing is inspired by Scottish Ceilidh d…
action, analysis, background, chemistry, dance, motion, pattern, videoAn Annotated Dataset For Near-Duplicate Detection In Personal Photo Collections Managing photo collections involves a variety of image quality assessmen…
copyright, detection, duplicate, groundtruth, retrievalThe Daimler Mono Pedestrian Detection Benchmark dataset contains a large training and test set. The training set contains 15.560 pedestrian samples (image…
detection, mono, object, outdoor, pedestrian, scale, urbanThe CALTECH 256 dataset by Li Fei-Fei contains 30607 images for 256 categories.
centered, classification, detection, image, object, sceneThe Airport MotionSeg dataset contains 12 sequences of videos of an aiprort scenario with small and large moving objects and various speeds. It is challen…
airport, camera, clustering, motion, segmentation, video, zoomThe CERTH image blur dataset consists of 2450 digital images, 1850 out of which are photographs captured by various camera models in different shooting co…
blur, defocus, detection, image, motion, qualityThe Fish4Knowledge project (groups.inf.ed.ac.uk/f4k/) is pleased to announce the availability of 2 subsets of our tropical coral reef fish video and ext…
animal, camera, classification, fish, motion, nature, recognition, video, waterThe YouTube-Objects dataset is composed of videos collected from YouTube by querying for the names of 10 object classes. It contains between 9 and 24 vide…
detection, flow, object, optical, segmentation, videoThe Aspect Layout dataset is designed to allow evaluation of object detection for aspect ratios in perspective images. Author text: In this project we…
aspect, detection, layout, object, perspective, ratioThe Longterm Pedestrian dataset consists of images from a stationary camera running 24 hours for 7 days at about 1 fps. It used for adaptive detection an…
background, change, coffee, detection, graz, illumination, indoor, multitarget, pedestrian, robustThe Microsoft COCO (mscoco) is an image recognition and segmentation dataset which contains more 300k images for more than 70 categories. Other features…
benchmark, context, detection, object, recognition, segmentation, semanticThe Leeds Cows dataset by Derek Magee consists of 14 different video sequences showing a total of 18 cows walking from right to left in front of different…
animal, background, cow, detection, segmentation, videoThe QMUL Junction dataset is a busy traffic scenario for research on activity analysis and behavior understanding. Video length: 1 hour (90000 frames)…
behavior, counting, crowd, detection, motion, pedestrian, tracking, videoThe dataset contains 15 documentary films that are downloaded from YouTube, whose durations vary from 9 minutes to as long as 50 minutes, and the total nu…
detection, object, videoMany different labeled video datasets have been collected over the past few years, but it is hard to compare them at a glance. So we have created a handy …
action, benchmark, classification, detection, object, recognition, videoThe TRaffic ANd COngestionS (TRANCOS) dataset, a novel benchmark for (extremely overlapping) vehicle counting in traffic congestion situations. It consist…
car, detection, highway, object, spain, traffic, transportation, urban, vehicleThe Multi-FoV synthetic datasets are two synthetic scenes (vehicle moving in a city, and flying robot hovering in a confined room). For each scene, three …
blender, camera, fov, groundtruth, odometry, synthetic, visualThe High Definition Analytics (HDA) dataset is a multi-camera High-Resolution image sequence dataset for research on High-Definition surveillance: Pedestr…
benchmark, camera, detection, high-definition, human, indoor, lisbon, multiview, network, pedestrian, re-identification, surveillance, tracking, videoThe PASCAL VOC is augmented with segmentation annotation for semantic parts of objects. For example, for the person category, we provide segmentation mask…
detection, human, object, part, pascal, pedestrian, recognition, segmentation, semanticThe Ford Car dataset is joint effort of Pandey et al. (for collecting images, Lidar points, calibration etc.) and us (for annotation of 2D and 3D objects)…
3d, car, detection, groundtruth, lidar, sfmThe mirror symmetry database contains 176 single-symmetry and 63 multyple-symmetry images (.png files) with accompanying ground-truth annotations (.mat fi…
detection, groundtruth, mirror, symmetryThe UMD Dynamic Scene Recognition dataset consists of 13 classes and 10 videos per class and is used to classify dynamic scenes. The dataset has been de…
classification, dynamic, motion, recognition, scene, videoGlobal Symmetry Ground-truth for AVA dataset Release Date: 2016 For detailed information, please refer to: Elawady, Mohamed, Ccile Barat, Christophe …
aesthetic, bilateral, detection, global, mirror, reflection, symmetryThe Where Who Why (WWW) dataset provides 10,000 videos with over 8 million frames from 8,257 diverse scenes, therefore offering a superior comprehensive d…
crowd, detection, flow, optical, pedestrian, recognition, surveillance, videoThe Quad 6K dataset is a Structure-from-Motion dataset taken at Arts Quad at Cornell University campus and consists of 6514 images with ground truth posit…
3d gps, 3d reconstruction, groundtruth, landmark, sfm, urbanThe SPHERE human skeleton movements dataset was created using a Kinect camera, that measures distances and provides a depth map of the scene instead of th…
action, behavior, depth, human, kinect, motion, movement, skeleton, videoThe Yotta dataset consists of 70 images for semantic labeling given in 11 classes. It also contains multiple videos and camera matrices for 14km or drivin…
3d, camera, classification, reconstruction, segmentation, semantic, urban, videoThe Daimler Urban Segmentation Dataset consists of video sequences recorded in urban traffic. The dataset consists of 5000 rectified stereo image pairs wi…
motion, outdoor, segmentation, semantic, stereo, urbanWIDER FACE dataset is a large-scale face detection benchmark dataset with 32,203 images and 393,703 face annotations, which have high degree of variabilit…
detection, face, occlusion, pose, scaleSVHN is a real-world image dataset for developing machine learning and object recognition algorithms with minimal requirement on data preprocessing and fo…
classification, detection, number, real, recognition, streetside, streetview, text, urban, worldThe contour patches dataset is a large dataset of images patch matches used for contour detection. References: C. L. Zitnick and D. Parikh The Role o…
contour, detection, edge, image, lowlevel, match, patch, segmentationA dataset acquired with 3 synchronized sensors (Primesense Carmine 1.09, Microsoft Kinect v2, Canon IXUS 950 IS), featuring: * 30 industry-relevant obje…
3d, estimation, object, pose, rgbd, texture-lessSince the publicly available face image datasets are often of small to medium size, rarely exceeding tens of thousands of images, and often without age in…
age, biometry, detection, face, imdb, recognition, wikipediaThe PIROPO database (People in Indoor ROoms with Perspective and Omnidirectional cameras) comprises multiple sequences recorded in two different indoor ro…
detection, fisheye, human, indoor, omnidirectional, people, perspective, room, surveillanceDataset A (former NLPR Gait Database) was created on Dec. 10, 2001, including 20 persons. Each person has 12 image sequences, 4 sequences for each of the …
action, biometry, classification, foot, gait, human, motion, pressure, recognitionThe Multi-illuminant Image Sequences dataset contains 16 video sequences (13 with single light source and 3 with two global light sources), recorded with…
balance, chromaticity, color, constancy, dichromatic, illumination, light, nature, object, physics, whiteYahoo Flickr Creative Commons 100M (YFCC100M) dataset contains a list of photos and videos. This list is compiled from data available on Yahoo! Flickr. Al…
3d, clustering, community, detection, flickr, image, internet, landmark, recognition, reconstruction, socialAt Udacity, we believe in democratizing education. How can we provide opportunity to everyone on the planet? We also believe in teaching really amazing an…
autonomous, car, classification, detection, driving, recognition, robot, segmentation, street, synthetic, time, urban, videoThe GaTech VideoStab dataset consists of N videos for the task of video stabilization. This code is implemented in Youtube video editor for stabilization.…
camera, path, stabilization, videoInstance recognition from depth data. Contains various challenges of Pose, Clutter, Occlusion and similar looking objects (Bonde, U., Badrinarayanan, V., …
depth, detection, instance, poseThe Berkeley Multimodal Human Action Database (MHAD) contains 11 actions performed by 7 male and 5 female subjects in the range 23-30 years of age except …
action, classification, motion, multiview, recognitionContains 6 object categories similar to object categories in Pascal VOC that are suitable for studying the abnormalities stemming from objects.
detectionThe CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test image…
color, image classification, object, patch, scene, tinyScanNet is an RGB-D video dataset containing 2.5 million views in more than 1500 scans, annotated with 3D camera poses, surface reconstructions, and insta…
3d, cad, indoor, layout, object, realism, recognition, rendering, room, scene, segmentation, synthetic30000+ frames with vehicle rear annotation and classification (car and trucks) on motorway/highway sequences. Annotation semi-automatically generated usin…
detectionThis UIUC Cars dataset by Shivani Agarwal, Aatif Awan and Dan Roth contains images of side views of cars for use in evaluating object detection algorithms…
car, detection, recognition, scale, sideview, urbanClassification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets
classification, detectionFor the first few decades of the fields existence, computer vision has been focused on algorithmic, logical approaches to perception. But it was only with…
3d, depth, indoor, kinect, object, recognition, reconstructionThe Farman Institute 3D Point Sets dataset contains 11 objects by a 3D laser scanner. This dataset was peer-reviewed by Image Processing On Line: Farman I…
3d, laser, model, object, point, reconstruction, scanner15 wide baseline stereo image pairs with large viewpoint change, provided ground truth homographies. Image size (~1000x700 pixels, RGB) D. Mishkin and…
description, detection, feature, matching, viewpoint, wide baseline stereoThis dataset contains 12,995 face images which are annotated with (1) five facial landmarks, (2) attributes of gender, smiling, wearing glasses, and head …
attribute, cnn, deep learning, detection, face, landmark detectionThe Swedish Traffic Sign Recognition provides Matlab code for parsing the annotation files and displaying the results. Part0 for each set contains the ann…
city, detection, recognition, sign, traffic, urbanThe .enpeda.. Image Sequence Analysis Test Site (EISATS) offers sets of long bi- or trinocular image sequences recorded in the context of vision-based dri…
analysis, flow, motion, optical, segmentation, semantic, stereo, visionThis material is supplementary to Michael Stark, Bernt Schiele. How Good are Local Features for Classes of Geometric Objects. Eleventh IEEE Internatio…
binary, classification, object, shape, toolA New Color Image Database for Benchmarking of Face Detection Techniques and Human Skin Segmentation Techniques. A new color face image database for be…
benchmarking, detection, face, segmentation, skinThe Daimler Mono Pedestrian Classification Benchmark dataset consists of two parts: a base data set. The base data set contains a total of 4000 pedestri…
classification, illumination, object, outdoor, pedestrian, scale, urbanJPL First-Person Interaction dataset (JPL-Interaction dataset) is composed of human activity videos taken from a first-person viewpoint. The dataset parti…
action, human, interactive, motion, recognition, videoThe Annotated Facial Landmarks in the Wild (AFLW) consists of a large-scale collection of annotated face images gathered from the web, exhibiting a large …
age, annotation, detection, face, landmark, poseThe Visual Attributes dataset contains visual attribute annotations for over 500 object classes (animate and inanimate) which are all represented in Image…
attribute, classification, imagenet, object, recognitionThis data set comprises 144 images of an edge profile cutting head of a milling machine. The head tool contains a total of 30 cutting inserts. The cutting…
cutting, edge, head, inserts, localization, milling, monitoring, object, profile, tool, tools, wearThe dataset consist of the about 50 hours obtained from kindergarten surveillance videos. Dataset, totally approximately 100 videos sequences (1000GB, 50 …
action, background, behavior, human, segmentation, video surveillanceMultispectral Imaging (MSI) datasets were acquired using IRIS II which is a lightweight portable system comprising of a high resolution camera, a novel fi…
alignment, groundtruth, illumination, matching, multi-spectral, registration, wavelengthThe CMP map2photo dataset consists of 6 pairs, where one image is satellite photo and second image is a map of the same area. The task is to match these …
baseline, description, detection, feature, map, matching, remote, sensing, wideThe FlickrLogos-32 dataset contains photos showing brand logos and is meant for the evaluation of multi-class logo recognition as well as logo retrieval m…
classification brand boundingbox, detection, flickr, image, logo, machine learning, object recognition, retrievalTh EPFL Multi-View Car dataset contains 20 sequences of cars as they rotate by 360 degrees. There is one image approximately every 3-4 degrees. Using the …
car, detection, estimation, multiview, pose, rotationThe Kendall Square webcam dataset consists of two streams for one sunny day and one cloudy day of a city square. It is used for tracking and analyzing col…
appearance, change, color, detection, sky, weather, webcamThe PETS 2016 IPATCH dataset contains a set of fourteen multi camera recordings (visible, themal) collected off the coast of Brest, France, in collaborati…
boat, detection, gps, maritime, multimodal, radar, thermal, tracking, vessel, visibleThe city planar and non-planar datset consists of urban scenes accompanied by text files describing the plane/non-plane locations. Training Set (Univer…
3d, building, detection, estimation, plane, urbanThe BEOID dataset includes object interactions ranging from preparing a coffee to operating a weight lifting machine and opening a door. The dataset is re…
3d, egocentric, interaction, object, pose, tracking, videoChairGest is an open challenge / benchmark. The task consists in spotting and recognizing gestures from multiple synchronized sensors: 1 Kinect and 4 Xse…
benchmark, detection, gesture, human, kinect, recognitionThe Stanford 40 Actions dataset contains images of humans performing 40 actions. In each image, we provide a bounding box of the person who is performing …
action, boundingbox, detection, human, recognitionThe Wide (multiple) Baseline Dataset. 31 image pairs, simultaneously combining several nuisance factors: geometry, illumination, IR-visible, etc. WxBS: …
day, description, detection, feature, ir, matching, night, viewpointThe TU Berlin Multi-Object and Multi-Camera Tracking Dataset (MOCAT) is a synthetic dataset to train and test tracking and detection systems in a virtual …
animal, detection, evaluation, multi-class, multi-view, pedestrian, synthetic, tracking, vehicleThe Oxford RobotCar Dataset contains over 100 repetitions of a consistent route through Oxford, UK, captured over a period of over a year. The dataset cap…
autonomous, car, classification, detection, driving, recognition, robot, segmentation, street, time, urban, video, yearCalifornia-ND contains 701 photos taken directly from a real user's personal photo collection, including many challenging non-identical near-duplicate cas…
detectionTo evaluate our method we designed a new ground truth database of 50 images. The following zip-files contain: Data, Segmentation, Labelling - Lasso, Label…
background, boundingbox, color, image segmentation, optimizationThe SUNCG dataset is a Large 3D Model Repository for Indoor Scenes. SUNCG is an ongoing effort to establish a richly-annotated, large-scale dataset of…
3d, indoor, layout, object, realism, recognition, rendering, room, scene, segmentation, syntheticWe wanted to have a collection of action recognition papers and results that everybody can use for reference. The site will work by the community principl…
action, benchmark, dataset, recognitionThe VOT2016 pixel-wise annotations dataset contains pixel-wise per-frame annotations for sequences from VOT2016 dataset. The annotation is in a form of BW…
annotation, mask, object, segmentation, tracking, visualThe Traffic Video dataset consists of X video of an overhead camera showing a street crossing with multiple traffic scenarios. The dataset can be downlo…
detection, overhead, road, tracking, traffic, urban, video, view10000 images of natural scenes grabbed on Flickr, with 2695 logos instances cut and pasted from the BelgaLogos dataset.
detectionThis is a subset of the dataset introduced in the SIGGRAPH Asia 2009 paper, Webcam Clip Art: Appearance and Illuminant Transfer from Time-lapse Sequences.…
camera, change, illumination, light, nature, static, time, urban, video, webcamThe Caltech Lanes dataset includes four clips taken around streets in Pasadena, CA at different times of day. The archive below includes 1225 individual…
caltech, detection, lane, pasadena, road, urbanThe Comprehensive Cars (CompCars) dataset contains data from two scenarios, including images from web-nature and surveillance-nature. The web-nature data …
attribute, car, classification, fine-grained, object, recognition, urban, vehicleThe Mall dataset was collected from a publicly accessible webcam for crowd counting and profiling research. Ground truth: Over 60,000 pedestrians were …
counting, crowd, detection, indoor, pedestrian, tracking, video, webcamThe HandNet dataset contains depth images of 10 participants hands non-rigidly deforming infront of a RealSense RGB-D camera. This dataset includes 214…
articulation, classification, detection, fingertip, hand, pose, rgbd, segmentation, videoWe present a new large-scale dataset that contains a diverse set of stereo video sequences recorded in street scenes from 50 different cities, with high q…
car, cities, detection, pedestrian, person, segmentation, semantic, stereo, urban, video, weaklyThe MSR Action datasets is a collection of various 3D datasets for action recognition. See details http://research.microsoft.com/en-us/um/people/zliu/a…
3d, action, detection, recognition, reconstruction, videoWe share our omnidirectional and panoramic image dataset (with annotations) to be used for human and car detection. Please reach through: http://cvrg.iyt…
car, detection, human, omnidirection, panorama, recognitionThe CALTECH 101 dataset by Li Fei-Fei contains images for 101 categories with about 40 to 800 images per category. Most categories have about 50 images at…
centered, image classification, natural-image, object, sceneThe crowd datasets are collected from a variety of sources, such as UCF and data-driven crowd datasets. The sequences are diverse, representing dense crow…
anomaly, crowd, detection, human, pedestrian, scene, understanding, videoThe Weather and Illumination Database (WILD) is an extensive database of high quality images of an outdoor urban scene, acquired every hour over all seaso…
camera, change, depth, estimation, illumination, light, newyork, static, time, urban, video, weather, webcamThese sequences were used for our video interpolation work described in High-quality video view interpolation using a layered representation, C.L. Zitn…
3d reconstruction, camera, depth, segmentationt is composed of food intake movements, recorded with Kinect V1 (320240 depth frame resolution), simulated by 35 volunteers for a total of 48 tests. The d…
age, behavior, food, groundtruth, human, intake, kinect, monitoring, pointcloud, trackingThe dataset captures 25 people preparing 2 mixed salads each and contains over 4h of annotated accelerometer and RGB-D video data. Annotated activities co…
action, activity, classification, detection, recognition, tracking, videoCollected in a clothing store. Captured with Kinect (640*480, about 30fps)
detection, trackingThis dataset contains 7 challenging volleyball activity classes annotated in 6 videos from professionals in the Austrian Volley League (season 2011/12). A…
action, activity recognition, analysis, detection, sport, video, volleyballWe introduce the Shelf dataset for multiple human pose estimation from multiple views. In addition we annotate the body joints in the Campus dataset from …
3d, capture, estimation, human, motion, multiple, pose, viewThe CVC Partial Occlusion Virtual Pedestrian datasets (CVC-01 to CVC-06) cover a range of scenarios of occluded pedestrians generated in a virtual and rea…
classification, detection, occlusion, pedestrian, synthetic, tracking, urbanThe 1DSfM Landmarks is a collection of community-based image reconstruction by Kyle Wilson and is comprised of 14 datasets with comparison to bundler grou…
3d, benchmark, city, groundtruth, landmark, reconstruction, urbanCOCO-Stuff augments the COCO dataset with pixel-level stuff annotations for 10,000 images. These annotations can be used for scene understanding tasks lik…
annotation, benchmark, captioning, coco, groundtruth, segmentation, semantic, stuff, thingsPhos is a color image database of 15 scenes captured under different illumination conditions. More particularly, every scene of the database contains 15 …
detectionThe video co-segmentation dataset contains 4 video sets which totally has 11 videos with 5 frames of each video labeled with the pixel-level ground-trut…
co-segmentation, dataset, segmentation, videoThe Stanford Dogs dataset contains images of 120 breeds of dogs from around the world. This dataset has been built using images and annotation from ImageN…
classification, detection, dogs, fine-grained categorizationThe Video Summarization (SumMe) dataset consists of 25 videos, each annotated with at least 15 human summaries (390 in total). The data consists of videos…
action, benchmark, event, groundtruth, human, summary, videoThe Graz02 dataset by Andreas Opelt and Axel Pinz contains four categories of images: bikes, people, cars and a single background class. The annotation ha…
background, bike, car, clutter, graz, object detection, pedestrianThe TUD Crossing dataset from Micha Andriluka, Stefan Roth and Bernt Schiele consists of 201 images with 1008 highly overlapping pedestrians with signific…
detection, multitarget, overlap, pedestrian, segmentation, sideview, tracking, urbanBelgiumTS is a large dataset with 10000+ traffic sign annotations, thousands of physically distinct traffic signs. 4 video sequences recorded with 8 high …
belgium, calibration, camera, classification, road, sign, traffic, urbanThe German Traffic Sign Recognition Benchmark is a dataset for multi-class detection problem in natural images and do cordially invite you to participate.…
detection, recognition, traffic, traffic sign, urban1521 images with human faces, recorded under natural conditions, i.e. varying illumination and complex background. The eye positions have been set manuall…
detection10000 images of natural scenes, with 37 different logos, and 2695 logos instances, annotated with a bounding box.
detectionThe PETS 2009 dataset contains 3 parts showing multi-view sequences containing pedestrians walking in an outdoor environment. The parts are used for perso…
detection, frontview, human, occlusion multitarget, outdoor, overlap, pedestrian, trackingThese datasets were generated for the M2CAI challenges, a satellite event of MICCAI 2016 in Athens. Two datasets are available for two different challenge…
challenge, medicine, recognition, surgery, video, workflowThe UK Bench dataset from Henrik Stewenius and David Nister contains 10200 images of N=2550 groups with each four images at size 640x480. The images are r…
centered, image retrieval, object, rotationWe collected a video dataset, termed ChokePoint, designed for experiments in person identification/verification under real-world surveillance conditions u…
clustering, detection, face, human, identification, multiview, pedestrian, real, recognition, sequence, surveillance, worldLabelMe is a web-based image annotation tool that allows researchers to label images and share the annotations with the rest of the community. If you use …
detectionThe test sequences provide interested researchers a real-world multi-view test data set captured in the blue-c portals. The data is meant to be used for t…
action, camera, multiview, segmentation, trackingThis dataset package contains the software and data used for Detection-based Object Labeling on the RGB-D Scenes Dataset as implemented in the paper: De…
3d, depth, indoor, kinect, object, recognition, reconstructionThe MOT Challenge is a framework for the fair evaluation of multiple people tracking algorithms. In this framework we provide: - A large collection of d…
3d, benchmark, benhttp://motchallenge.net/chmark, dataset, evaluation, multiple, pedestrian, people, surveillance, target, tracking, videoThe FaceScrub dataset comprises a total of 107818 unconstrained face images of 530 celebrities crawled from the Internet, with about 200 images per person…
celebrity, detection, face, human, people, recognitionThe Extreme Zoom Dataset. EZD is a 6 image sets with incleasing zoom factor from general scene view to focusing on single detail. MODS: Fast and Robust …
description, detection, feature, matching, viewpoint, zoomThe Multicamera Human Action Video Data (MuHAVi) Manually Annotated Silhouette Data (MAS) are two datasets consisting of selected action sequences for th…
action, background, behavior, human, segmentation, video15,560 pedestrian and non-pedestrian samples (image cut-outs) and 6744 additional full images not containing pedestrians for bootstrapping. The test set c…
detectionWe present the 2017 DAVIS Challenge, a public competition specifically designed for the task of video object segmentation. Following the footsteps of othe…
benchmark, code, hd, object, quality, resolution, segmentation, tracking, video segmentationThe UrbanStreet dataset used in the paper can be downloaded here [188M] . It contains 18 stereo sequences of pedestrians taken from a stereo rig mounted o…
detection, human, multitarget, pedestrian, recognition, segmentation, tracking, urban, videoThe Inria Aerial Image Labeling addresses a core topic in remote sensing: the automatic pixelwise labeling of aerial imagery (link to paper). Dataset fe…
aerial, building, city, footprint, groundtruth, house, segmentation, semantic, urbanA 66 stereo pairs dataset with their subpixel ground truths. The construction and improvement of algorithms for subpixel stereovision requires very prec…
3d, depth, groundtruth, noise, pointcloud, stereo, stereovision, subpixelContains drawing pages from US patents with manually labeled figure and part labels.
detectionThe Graz01 dataset by Andreas Opelt and Axel Pinz contains four types of images: bikes, people, background with no bikes, background with no people.
background, bike, clutter, graz, object detection, occlusion, pedestrian