<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:academictorrents="http://academictorrents.com/" version="2.0">
<channel>
<title>Computer Vision - Academic Torrents</title>
<description>collection curated by joecohen</description>
<link>https://academictorrents.com/collection/computer-vision</link>
<item>
<title>CelebV-HQ (Dataset)</title>
<description>Large-scale datasets have played indispensable roles in the recent success of face generation/editing and significantly facilitated the advances of emerging research fields. However, the academic community still lacks a video dataset with diverse facial attribute annotations, which is crucial for the research on face-related videos. In this work, we propose a large-scale, high-quality, and diverse video dataset with rich facial attribute annotations, named the High-Quality Celebrity Video Dataset (CelebV-HQ). CelebV-HQ contains 35,666 video clips with the resolution of 512x512 at least, involving 15,653 identities. All clips are labeled manually with 83 facial attributes, covering appearance, action, and emotion. We conduct a comprehensive analysis in terms of age, ethnicity, brightness stability, motion smoothness, head pose diversity, and data quality to demonstrate the diversity and temporal coherence of CelebV-HQ. Besides, its versatility and potential are validated on two representative tasks, i.e., unconditional video generation and video facial attribute editing. Furthermore, we envision the future potential of CelebV-HQ, as well as the new opportunities and challenges it would bring to related research directions.</description>
<link>https://academictorrents.com/download/843b5adb0358124d388c4e9836654c246b988ff4</link>
</item>
<item>
<title>Amos: A large-scale abdominal multi-organ benchmark for versatile medical image segmentation (Dataset)</title>
<description>AMOS provides 500 CT and 100 MRI scans collected from multi-center, multi-vendor, multi-modality, multi-phase, multi-disease patients, each with voxel-level annotations of 15 abdominal organs, providing challenging examples and test-bed for studying robust segmentation algorithms under diverse targets and scenarios. We further benchmark several state-of-the-art medical segmentation models to evaluate the status of the existing methods on this new challenging dataset. We have made our datasets, benchmark servers, and baselines publicly available, and hope to inspire future research. </description>
<link>https://academictorrents.com/download/8277ce3d862883f08846d87099e3af4d89fd94c1</link>
</item>
<item>
<title>TotalSegmentator CT Dataset V2 (Dataset)</title>
<description>Info: This is version 2 of the TotalSegmentator dataset. In 1228 CT images we segmented 117 anatomical structures covering a majority of relevant classes for most use cases. The CT images were randomly sampled from clinical routine, thus representing a real world dataset which generalizes to clinical application. The dataset contains a wide range of different pathologies, scanners, sequences and institutions. </description>
<link>https://academictorrents.com/download/1dfeb3186514b40a2c212c21d494c665766bfbf4</link>
</item>
<item>
<title>Sculptures 6k dataset (Dataset)</title>
<description>The Sculptures 6k Dataset consists of 6340 images images collected from Flickr by searching for sculptures by Henry Moore and Auguste Rodin.  The dataset is split equally into a train and test set, each containing 3170 images.  For each set 10 different Henry Moore sculptures are chosen as query objects, and for each of these objects 7 images and query regions are defined, thus providing 70 queries for performance evaluation purposes.</description>
<link>https://academictorrents.com/download/45e3d613c72dbd3f7e0a5b27d1955b67ace0a014</link>
</item>
<item>
<title>SegThy Open-Access Dataset for Thyroid and Neck Segmentation (Dataset)</title>
<description>## Motivation Ultrasound (US) imaging plays a central role in the diagnosis of thyroid diseases as well as different pathologies of the neck region. Additionally, US is used for treatment planning, in the case of radioiodine therapy, and also as mean of following up on the success of different therapeutic efforts. Yet, in the vast majority of hospitals, 2D freehand US is applied. This type of examination has shown to have a high intra-observer and inter-observer variability [1], and low accuracy in terms of volume prediction in a thyroid volumetry setting using MRI as ground truth. To tackle these problems, we propose the use of 3D US in combination with machine learning (ML) algorithms to automatically segment the most relevant organs in the region, as well as anomalies, such as nodules and tumors. 3D US can be acquired in several ways, e.g. using 3D US probes, robotic US, so-called wobbler US probes, or using freehand tracked US. The last option has gained traction over the last years as it enables to extend almost any existing commercial 2D US at a low cost. On the side of ML, its availability and increasing computing power are flooding the medical world. Yet, ML can only provide trustworthy results if sufficient (annotated) data is available to “learn” from it. This motivated us to create and publish this dataset, and thus make a relevant contribution to the community. ## Citation [1] Tracked 3D ultrasound and deep neural network-based thyroid segmentation reduce interobserver variability in thyroid volumetry; M. Krönke, C. Eilers, D. Dimova, M. Köhler, G. Buschner, L. Schweiger, L. Konstantinidou, M. Makowski, J. Nagarajah, N. Navab, W. Weber, and T. Wendler; PLOS ONE - July 29, 2022 -  ## Dataset Description Sub-dataset 1 This sub-dataset consists of 28 (healthy) volunteers who were scanned using freehand tracked ultrasonography (both sides of their neck, focus on thyroid). A Siemens Acuson NX-3 US machine, combined with a 12MHz VF12-4 probe, was adapted to be tracked using electromagnetic tracking using the Piur tUS system. Additionally, all patients were imaged with MRI (Siemens Biograph mMR) using a T1-weighted VIBE (volumetric interpolated breath-hold) sequence centered in the thyroid area of the neck. The magnetic field was set to 3T. In practical terms, each volunteer in this subdataset has: 1x MRI (T1-weighted VIBE sequence) of the neck area, with voxel size 0.625x0.625x1 mm3 and field of view 320x320x80 mm3. 9x two-sided 3D US covering the thyroid region with voxel size of 0.12x0.12x0.12 mm3 and variable field of view. The nine 3D US were acquired by three physicians (three scans each). Label maps for the thyroid in all MR and US images. Sub-dataset 2 The second subdataset consists of 186 patients undergoing routinary thyroid US either as initial diagnostics means or as follow-up. The same US device and freehand tracked US extension device were used. Each patient in this sub-dataset has 6x two-sided 3D US covering the thyroid. The 6 3D US were acquired by 2 physicians (three scans each). In both dataset, the age and the sex of the patients is included in the metadata. The dataset is currently being extended to include the labels of the trachea, the common carotid arteries, the internal jugular veins, and (if present) thyroid nodules in sub-dataset 1 (in both US and MRI), and of the same organs in sub-dataset 2 in US. See figures as examples for the multilabel annotations. ## License This dataset is licensed under a CC BY license. This means that users of it can distribute, shuffle, adapt, and build upon the material with the sole condition that the creators are acknowledged. The license permits commercial use as long as the authors are cited.</description>
<link>https://academictorrents.com/download/a6530eb901e8c1c127166d1bebeffb0129f5bf9f</link>
</item>
<item>
<title>STructured Analysis of the Retina (Dataset)</title>
<description>The STARE (STructured Analysis of the Retina) Project was conceived and initiated in 1975 by Michael Goldbaum, M.D., at the University of California, San Diego. It was funded by the U.S. National Institutes of Health . During its history, over thirty people contributed to the project, with backgrounds ranging from medicine to science to engineering. Images and clinical data were provided by the Shiley Eye Center at the University of California, San Diego, and by the Veterans Administration Medical Center in San Diego. I had the pleasure of working on the project from 1996-2004. The contents of this web page reflect my contributions. Please contact me if you have any questions or requests concerning our data or code. Please contact Dr. Goldbaum if you have any requests concerning the current state of the project. # A brief overview of the project An ophthalmologist is a medical doctor that specializes in the structure, function, and diseases of the human eye. During a clinical examination, an opthalmologist notes findings that are visible in the eyes of the subject. The ophthalmologist then uses these findings to reason about the health of the subject. For instance, a patient may exhibit discoloration of the optic nerve, or a narrowing of the blood vessels in the retina. An opthalmologist uses this information to diagnose the patient, as having for instance Coats  disease or a central retinal artery occlusion. A common procedure during an examination is retinal imaging. An optical camera is used to see through the pupil of the eye to the rear inner surface of the eyeball. A picture is taken showing the optic nerve, fovea, surrounding vessels, and the retinal layer. The opthalmologist can then reference this image while considering any observed findings. This research concerns a system to automatically diagnose diseases of the human eye. The system takes as input information observable in a retinal image. This information is formulated to mimic the findings that an ophthalmologist would note during a clinical examination. The main output of the system is a diagnosis formulated to mimic the conclusion that an ophthalmologist would reach about the health of the subject. Our approach breaks the problem into two components. The first component concerns automatically processing a retinal image to denote the important findings. The second component concerns automatically reasoning about the findings to determine a diagnosis. Additional outputs include detailed measurements of the anatomical structures and lesions visible in the retinal image. These measurements are useful for tracking disease severity and the evaluation of treatment progress over time. By collecting a database of measurements for a large number of people, the STARE project could support clinical population studies and intern training.  # Papers A lot has been published on this project by many people; these are my two most relevant papers: A. Hoover, V. Kouznetsova and M. Goldbaum, "Locating Blood Vessels in Retinal Images by Piece-wise Threhsold Probing of a Matched Filter Response", IEEE Transactions on Medical Imaging , vol. 19 no. 3, pp. 203-210, March 2000. A. Hoover and M. Goldbaum, "Locating the optic nerve in a retinal image using the fuzzy convergence of the blood vessels", IEEE Transactions on Medical Imaging , vol. 22 no. 8, pp. 951-958, August 2003.</description>
<link>https://academictorrents.com/download/e4554cd63400dc13b74477efe98032c10757c269</link>
</item>
<item>
<title>TotalSegmentator CT Dataset (Dataset)</title>
<description> In 1204 CT images we segmented 104 anatomical structures (27 organs, 59 bones, 10 muscles, 8 vessels) covering a majority of relevant classes for most use cases. The CT images were randomly sampled from clinical routine, thus representing a real world dataset which generalizes to clinical application. The dataset contains a wide range of different pathologies, scanners, sequences and institutions.     s0720/segmentations/portal_vein_and_splenic_vein.nii.gz187.74kB s0720/segmentations/pancreas.nii.gz45.25kB s0720/segmentations/lung_upper_lobe_right.nii.gz218.92kB s0720/segmentations/lung_upper_lobe_left.nii.gz230.82kB s0720/segmentations/lung_middle_lobe_right.nii.gz201.18kB s0720/segmentations/lung_lower_lobe_right.nii.gz240.63kB s0720/segmentations/lung_lower_lobe_left.nii.gz239.49kB s0720/segmentations/liver.nii.gz273.08kB s0720/segmentations/kidney_right.nii.gz198.91kB s0720/segmentations/kidney_left.nii.gz197.82kB s0720/segmentations/inferior_vena_cava.nii.gz48.43kB s0720/segmentations/iliopsoas_right.nii.gz59.12kB s0720/segmentations/iliopsoas_left.nii.gz59.75kB s0720/segmentations/iliac_vena_right.nii.gz188.90kB s0720/segmentations/iliac_vena_left.nii.gz189.66kB s0720/segmentations/iliac_artery_right.nii.gz186.75kB s0720/segmentations/iliac_artery_left.nii.gz186.60kB s0720/segmentations/humerus_right.nii.gz41.96kB s0720/segmentations/humerus_left.nii.gz43.13kB s0720/segmentations/hip_right.nii.gz223.33kB s0720/segmentations/hip_left.nii.gz223.06kB s0720/segmentations/heart_ventricle_right.nii.gz48.07kB s0720/segmentations/heart_ventricle_left.nii.gz45.10kB s0720/segmentations/heart_myocardium.nii.gz49.17kB s0720/segmentations/heart_atrium_right.nii.gz44.41kB s0720/segmentations/heart_atrium_left.nii.gz43.02kB s0720/segmentations/gluteus_minimus_right.nii.gz46.65kB s0720/segmentations/gluteus_minimus_left.nii.gz45.95kB s0720/segmentations/gluteus_medius_right.nii.gz53.75kB s0720/segmentations/gluteus_medius_left.nii.gz52.68kB s0720/segmentations/gluteus_maximus_right.nii.gz58.02kB s0720/segmentations/gluteus_maximus_left.nii.gz56.20kB s0720/segmentations/gallbladder.nii.gz42.20kB s0720/segmentations/femur_right.nii.gz192.93kB s0720/segmentations/femur_left.nii.gz193.47kB s0720/segmentations/face.nii.gz183.15kB s0720/segmentations/esophagus.nii.gz188.93kB s0720/segmentations/duodenum.nii.gz189.53kB s0720/segmentations/colon.nii.gz239.38kB s0720/segmentations/clavicula_right.nii.gz42.92kB s0720/segmentations/clavicula_left.nii.gz42.50kB s0720/segmentations/brain.nii.gz183.15kB s0720/segmentations/autochthon_right.nii.gz62.97kB s0720/segmentations/autochthon_left.nii.gz63.75kB s0720/segmentations/aorta.nii.gz202.39kB s0720/segmentations/adrenal_gland_right.nii.gz184.50kB s0720/segmentations/adrenal_gland_left.nii.gz184.35kB s0720/ct.nii.gz      </description>
<link>https://academictorrents.com/download/337819f0e83a1c1ac1b7262385609dad5d485abf</link>
</item>
<item>
<title>COCO 2017 Resized to 256x256 (Dataset)</title>
<description>COCO: Common Objects in Context Resized to 256x245</description>
<link>https://academictorrents.com/download/eea5a532dd69de7ff93d5d9c579eac55a41cb700</link>
</item>
<item>
<title>Breast Ultrasound Images Dataset (Dataset BUSI) (Dataset)</title>
<description>The data collected at baseline include breast ultrasound images among women in ages between 25 and 75 years old. This data was collected in 2018. The number of patients is 600 female patients. The dataset consists of 780 images with an average image size of 500*500 pixels. The images are in PNG format. The ground truth images are presented with original images. The images are categorized into three classes, which are normal, benign, and malignant. If you use this dataset, please cite: Al-Dhabyani W, Gomaa M, Khaled H, Fahmy A. Dataset of breast ultrasound images. Data in Brief. 2020 Feb;28:104863. DOI: 10.1016/j.dib.2019.104863. | Subject area               | Medicine and Dentistry                                                                                                                                                             | |&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;|&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;&amp;mdash;| | More specific subject area | Radiology and Imaging                                                                                                                                                              | | Type of data               | Images and mask images                                                                                                                                                             | | How data was acquired      | LOGIQ E9 ultrasound and LOGIQ E9 Agile ultrasound system                                                                                                                           | | Data format                | PNG                                                                                                                                                                                | | Experimental factors       | All images are classified as normal, benign and malignant                                                                                                                          | | Experimental features      | When medical images are used for training deep learning models, they provide fast and accurate results in classification, detection, and segmentation of breast cancer.            | | Data source location       | Baheya Hospital for Early Detection &amp; Treatment of Women s Cancer, Cairo, Egypt.                                                                                                   |</description>
<link>https://academictorrents.com/download/d0b7b7ae40610bbeaea385aeb51658f527c86a16</link>
</item>
<item>
<title>UT Zappos50K (Version 2.1) (Dataset)</title>
<description>UT Zappos50K (UT-Zap50K) is a large shoe dataset consisting of 50,025 catalog images collected from Zappos.com. The images are divided into 4 major categories — shoes, sandals, slippers, and boots — followed by functional types and individual brands. The shoes are centered on a white background and pictured in the same orientation for convenient analysis. This dataset is created in the context of an online shopping task, where users pay special attentions to fine-grained visual differences. For instance, it is more likely that a shopper is deciding between two pairs of similar men s running shoes instead of between a woman s high heel and a man s slipper. GIST and LAB color features are provided. In addition, each image has 8 associated meta-data (gender, materials, etc.) labels that are used to filter the shoes on Zappos.com.  # Citation This dataset is for academic, non-commercial use only. If you use this dataset in a publication, please cite the following papers: A. Yu and K. Grauman. "Fine-Grained Visual Comparisons with Local Learning". In CVPR, 2014. [paper] [supp] [poster] [bibtex] [project page]</description>
<link>https://academictorrents.com/download/3b3cb58f4ccafc6320d06d00f0862a4ba923b510</link>
</item>
<item>
<title>Whale Shark ID Dataset (Dataset)</title>
<description>Our released whale shark (Rhincodon typus) data set represents a collaborative effort based on the data collection and population modeling efforts conducted at Ningaloo Marine Park in Western Australia from 1995-2008 (Holmberg et al. 2008, 2009). Photos (7888) and metadata from 2441 whale shark encounters were collected from 464 individual contributors, especially from the original research of Brad Norman and from members of the local whale shark tourism industry who sight these animals annually from April-June. Images were annotated with bounding boxes around each visible whale shark and viewpoints labeled (e.g., left, right, etc.). A total of 543 individual whale sharks were identified by their unique spot patterning using first computer-assisted spot pattern recognition (Arzoumanian et al. 2005) and then manual review and confirmation.  A total of 7,693 named sightings were exported. The dataset is released in the Microsoft COCO format () and therefore uses flat image folders with associated YAML metadata files. We have collapsed the entire dataset into a single "train" label and have left "val" and "test" empty; we do this as an invitation to researchers to experiment with their own novel approaches for dealing with the unbalanced and chaotic distribution on the number of sightings per individual.  All of the images in the dataset have been resized to have a maximum linear dimension of 3,000 pixels.  The metadata for all animal sightings is defined by an axis-aligned bounding box via and includes information on the rotation of the box (theta), the viewpoint of the animal, a species (category) ID, a source image ID, an individual string ID name, and other miscellaneous values.  The temporal ordering of the images, and an anonymized ID for the original photographer, can be determined from the metadata for each image. For research or press contact, please direct all correspondence to Wild Me at info@wildme.org.  Wild Me () is a registered 501(c)(3) not-for-profit based in Portland, Oregon, USA and brings state-of-the-art computer vision tools to ecology researchers working around the globe on wildlife conservation. Direct download mirror: </description>
<link>https://academictorrents.com/download/bb47cd1d6dde2f49b040495382c778c102409080</link>
</item>
<item>
<title>Great Zebra and Giraffe Count ID Dataset (Dataset)</title>
<description>Our dataset for plains zebra (Equus quagga) is taken from a two-day census of the Nairobi National Park, located just south of the capital’s airport in Nairobi, Kenya.  The “Great Zebra and Giraffe Count” (GZGC) photographic census was organized on February 28th and March 1st 2015 and had the participation of 27 different teams of citizen scientists, 55 total photographers, and collected 9,406 images of plains zebra and Masai giraffe (Giraffa tippelskirchi) (Parham et al. 2017).  Only images containing either zebras or giraffes were included in the exported dataset, a total of 4,948 images, where the original biographical information of the original contributors are removed.  All images are labeled with bounding boxes around the individual animals for which there is ID metadata, meaning some images contain missing boxes and are not intended to be used for object detection training or testing.  Viewpoints for all animal annotations were also added.  All ID assignments were completed using the HotSpotter algorithm (Crall et al. 2013) by visually matching the stripes and spots as seen on the body of the animal.  A total of 2,056 combined names are released for 6,286 individual zebra and 639 giraffe sightings.  This dataset presents as a challenging comparison compared to the whale shark dataset since it contains a significantly higher number of animals that are only seen once during the survey. The dataset is released in the Microsoft COCO format () and therefore uses flat image folders with associated YAML metadata files. We have collapsed the entire dataset into a single "train" label and have left "val" and "test" empty; we do this as an invitation to researchers to experiment with their own novel approaches for dealing with the unbalanced and chaotic distribution on the number of sightings per individual.  All of the images in the dataset have been resized to have a maximum linear dimension of 3,000 pixels.  The metadata for all animal sightings is defined by an axis-aligned bounding box via and includes information on the rotation of the box (theta), the viewpoint of the animal, a species (category) ID, a source image ID, an individual string ID name, and other miscellaneous values.  The temporal ordering of the images, and an anonymized ID for the original photographer, can be determined from the metadata for each image. For research or press contact, please direct all correspondence to Wild Me at info@wildme.org.  Wild Me () is a registered 501(c)(3) not-for-profit based in Portland, Oregon, USA and brings state-of-the-art computer vision tools to ecology researchers working around the globe on wildlife conservation. Direct download mirror: </description>
<link>https://academictorrents.com/download/69160c6bf11275321017f18124dbaff2d381b21c</link>
</item>
<item>
<title>Object-CXR - Automatic detection of foreign objects on chest X-rays (Dataset)</title>
<description>## Data 5000 frontal chest X-ray images with foreign objects presented and 5000 frontal chest X-ray images without foreign objects were filmed and collected from about 300 township hosiptials in China. 12 medically-trained radiologists with 1 to 3 years of experience annotated all the images. Each annotator manually annotates the potential foreign objects on a given chest X-ray presented within the lung field. Foreign objects were annotated with bounding boxes, bounding ellipses or masks depending on the shape of the objects. Support devices were excluded from annotation. A typical frontal chest X-ray with foreign objects annotated looks like this:  ## Annotation Object-level annotations for each image, which indicate the rough location of each foreign object using a closed shape. Annotations are provided in csv files and a csv example is shown below.    csv image_path,annotation /path/#####.jpg,ANNO_TYPE_IDX x1 y1 x2 y2;ANNO_TYPE_IDX x1 y1 x2 y2 ... xn yn;... /path/#####.jpg, /path/#####.jpg,ANNO_TYPE_IDX x1 y1 x2 y2 ...     Three type of shapes are used namely rectangle, ellipse and polygon. We use  0 ,  1  and  2  as  ANNO_TYPE_IDX  respectively. - For rectangle and ellipse annotations, we provide the bounding box (upper left and lower right) coordinates in the format  x1 y1 x2 y2  where  x1  &lt;  x2  and  y1  &lt;  y2 . - For polygon annotations, we provide a sequence of coordinates in the format  x1 y1 x2 y2 ... xn yn . &gt; ### Note: &gt; Our annotations use a Cartesian pixel coordinate system, with the origin (0,0) in the upper left corner. The x coordinate extends from left to right; the y coordinate extends downward. ## Organizers [JF Healthcare]() is the primary organizer of this challenge.</description>
<link>https://academictorrents.com/download/fdc91f11d7010f7259a05403fc9d00079a09f5d5</link>
</item>
<item>
<title>SIIM-ACR Pneumothorax Segmentation (Dataset)</title>
<description>In this competition, you’ll develop a model to classify (and if present, segment) pneumothorax from a set of chest radiographic images. If successful, you could aid in the early recognition of pneumothoraces and save lives. What am I predicting? We are attempting to a) predict the existence of pneumothorax in our test images and b) indicate the location and extent of the condition using masks. Your model should create binary masks and encode them using RLE. </description>
<link>https://academictorrents.com/download/6ef7c6d039e85152c4d0f31d83fa70edc4aba088</link>
</item>
<item>
<title>Leaf counting dataset (Dataset)</title>
<description>## Leaf counting dataset Dataset containing  9372 RGB images of weeds with the number of leaves counted. The images are collected in fields across Denmark using Nokia and Samsung cell phone cameras; Samsung, Nikon, Canon and Sony consumer cameras; and a Point Grey industrial camera.  ## Citation If you use this dataset in your research or elsewhere, please cite/reference the following paper: PAPER: Weed Growth Stage Estimator Using Deep Convolutional Neural Networks Bibtex</description>
<link>https://academictorrents.com/download/a147c27ea0a9c155df9d77af832c321210cf5529</link>
</item>
<item>
<title>MPII Human Pose Dataset (Dataset)</title>
<description>MPII Human Pose dataset is a state of the art benchmark for evaluation of articulated human pose estimation. The dataset includes around 25K images containing over 40K people with annotated body joints. The images were systematically collected using an established taxonomy of every day human activities. Overall the dataset covers 410 human activities and each image is provided with an activity label. Each image was extracted from a YouTube video and provided with preceding and following un-annotated frames. In addition, for the test set we obtained richer annotations including body part occlusions and 3D torso and head orientations. Following the best practices for the performance evaluation benchmarks in the literature we withhold the test annotations to prevent overfitting and tuning on the test set. We are working on an automatic evaluation server and performance analysis tools based on rich test set annotations. Citing the dataset</description>
<link>https://academictorrents.com/download/6be335f0d038fd4ed4422dd318705e0843059718</link>
</item>
<item>
<title>LNDb CT scan dataset (training) (Dataset)</title>
<description>The main goal of this challenge is the automatic classification of chest CT scans according to the 2017 Fleischner society pulmonary nodule guidelines for patient follow-up recommendation. The LNDb dataset contains 294 CT scans collected retrospectively at the Centro Hospitalar e Universitário de São João (CHUSJ) in Porto, Portugal between 2016 and 2018. All data was acquired under approval from the CHUSJ Ethical Commitee and was anonymised prior to any analysis to remove personal information except for patient birth year and gender. Further details on patient selection and data acquisition can be consulted on the database description paper. Each CT scan was read by at least one radiologist at CHUSJ to identify pulmonary nodules and other suspicious lesions. A total of 5 radiologists with at least 4 years of experience reading up to 30 CTs per week participated in the annotation process throughout the project. Annotations were performed in a single blinded fashion, i.e. a radiologist would read the scan once and no consensus or review between the radiologists was performed. Each scan was read by at least one radiologist. The instructions for manual annotation were adapted from LIDC-IDRI. Each radiologist identified the following lesions: - nodule ⩾3mm: any lesion considered to be a nodule by the radiologist with greatest in-plane dimension larger or equal to 3mm; - nodule &lt;3mm: any lesion considered to be a nodule by the radiologist with greatest in-plane dimension smaller than 3mm; - non-nodule: any pulmonary lesion considered not to be a nodule by the radiologist, but that contains features which could make it identifiable as a nodule; The annotation process varied for the different categories. Nodules ⩾3mm were segmented and subjectively characterized according to LIDC-IDRI (ratings on subtlety, internal structure, calcification, sphericity, margin, lobulation, spiculation, texture and likelihood of malignancy). For a complete description of these characteristics the reader is referred to McNitt-Gray et al.. For nodules &lt;3mm the nodule centroid was marked and subjective assessment of the nodule s characteristics was performed. For non-nodules, only the lesion centroid was marked. Given that different radiologists may have read the same CT and no consensus review was performed, variability in radiologist annotations is expected. Note that from the 294 CTs of the LNDb dataset, 58 CTs with annotations by at least two radiologists have been withheld for the test set, as well as the corresponding annotations. </description>
<link>https://academictorrents.com/download/e3c196b07c8ea94ac5fca872bccf2cc035f4e88d</link>
</item>
<item>
<title>Illinois DOC labeled faces dataset (Dataset)</title>
<description>This is a dataset of prisoner mugshots and associated data (height, weight, etc). The copyright status is public domain, since it s produced by the government, the photographs do not have sufficient artistic merit, and a mere collection of facts aren t copyrightable. The source is the Illinois Dept. of Corrections. In total, there are 68149 entries, of which a few hundred have shoddy data. It s useful for neural network training, since it has pictures from both front and side, and they re (manually) labeled with date of birth, name (useful for clustering), weight, height, hair color, eye color, sex, race, and some various goodies such as sentence duration and whether they re sex offenders. Here is the readme file: &amp;mdash;-BEGIN README&amp;mdash;- Scraped from the Illinois DOC.</description>
<link>https://academictorrents.com/download/4b9b7e449aa732842aea1a7d4e6413f4507aea99</link>
</item>
<item>
<title>NIH Chest X-ray Dataset (Resized to 224x224) (Dataset)</title>
<description>This dataset is resized versions of images to 224x224. ![]() (1, Atelectasis; 2, Cardiomegaly; 3, Effusion; 4, Infiltration; 5, Mass; 6, Nodule; 7, Pneumonia; 8, Pneumothorax; 9, Consolidation; 10, Edema; 11, Emphysema; 12, Fibrosis; 13, Pleural_Thickening; 14 Hernia) ### Background &amp; Motivation: Chest X-ray exam is one of the most frequent and cost-effective medical imaging examination. However clinical diagnosis of chest X-ray can be challenging, and sometimes believed to be harder than diagnosis via chest CT imaging. Even some promising work have been reported in the past, and especially in recent deep learning work on Tuberculosis (TB) classification. To achieve clinically relevant computer-aided detection and diagnosis (CAD) in real world medical sites on all data settings of chest X-rays is still very difficult, if not impossible when only several thousands of images are employed for study. This is evident from [2] where the performance deep neural networks for thorax disease recognition is severely limited by the availability of only 4143 frontal view images [3] (Openi is the previous largest publicly available chest X-ray dataset to date). In this database, we provide an enhanced version (with 6 more disease categories and more images as well) of the dataset used in the recent work [1] which is approximately 27 times of the number of frontal chest x-ray images in [3]. Our dataset is extracted from the clinical PACS database at National Institutes of Health Clinical Center and consists of ~60% of all frontal chest x-rays in the hospital. Therefore we expect this dataset is significantly more representative to the real patient population distributions and realistic clinical diagnosis challenges, than any previous chest x-ray datasets. Of course, the size of our dataset, in terms of the total numbers of images and thorax disease frequencies, would better facilitate deep neural network training [2]. Refer to [1] on the details of how the dataset is extracted and image labels are mined through natural language processing (NLP). ### Details: ChestX-ray dataset comprises 112,120 frontal-view X-ray images of 30,805 unique patients with the text-mined fourteen disease image labels (where each image can have multi-labels), mined from the associated radiological reports using natural language processing. Fourteen common thoracic pathologies include Atelectasis, Consolidation, Infiltration, Pneumothorax, Edema, Emphysema, Fibrosis, Effusion, Pneumonia, Pleural_thickening, Cardiomegaly, Nodule, Mass and Hernia, which is an extension of the 8 common disease patterns listed in our CVPR2017 paper. Note that original radiology reports (associated with these chest x-ray studies) are not meant to be publicly shared for many reasons. The text-mined disease labels are expected to have accuracy &gt;90%.Please find more details and benchmark performance of trained models based on 14 disease labels in our arxiv paper:  ### Contents: 1. 112,120 frontal-view chest X-ray PNG images in 1024*1024 resolution (under images folder) 2. Meta data for all images (Data_Entry_2017.csv): Image Index, Finding Labels, Follow-up #, Patient ID, Patient Age, Patient Gender, View Position, Original Image Size and Original Image Pixel Spacing. 3. Bounding boxes for ~1000 images (BBox_List_2017.csv):Image Index, Finding Label, Bbox[x, y, w, h]. [x y] are coordinates of each box s topleft corner. [w h] represent the width and height of each box. If you find the dataset useful for your research projects, please cite our CVPR 2017 paper:Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, MohammadhadiBagheri, Ronald M. Summers.ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases, IEEE CVPR, pp. 3462-3471,2017</description>
<link>https://academictorrents.com/download/e615d3aebce373f1dc8bd9d11064da55bdadede0</link>
</item>
<item>
<title>Ocular Disease Intelligent Recognition ODIR-5K (Dataset)</title>
<description>We collected a structured ophthalmic database of 5,000 patients with age, color fundus photographs from left and right eyes and doctors  diagnostic keywords from doctors (in short, ODIR-5K). This dataset is ‘‘real-life’’ set of patient information collected by Shanggong Medical Technology Co., Ltd. from different hospitals/medical centers in China. In these institutions, fundus images are captured by various cameras in the market, such as Canon, Zeiss and Kowa, resulting into varied image resolutions. Patient identifying information will be removed. Annotations are labeled by trained human readers with quality control management. They classify patient into eight labels including normal (N), diabetes (D), glaucoma (G), cataract (C), AMD (A), hypertension (H), myopia (M) and other diseases/abnormalities (O) based on both eye images and additionally patient age. The publishing of this dataset follows the ethical and privacy rules of China. Table 1 shows one record from ODIR-5K dataset. The 5,000 patients in this challenge are divided into training, off-site testing and on-site testing subsets. Almost 4,000 cases are used in training stage while others are for testing stages (off-site and on-site). Table 2 shows the distribution of case number with respect to eight labels in different stages. Note: one patient may contains one or multiple labels.  </description>
<link>https://academictorrents.com/download/cf3b8d5ecdd4284eb9b3a80fcfe9b1d621548f72</link>
</item>
<item>
<title>Minecraft Skins (Dataset)</title>
<description>An image data set containing 900,000+ Images of unique Minecraft skins of real players. Could be used for training a GAN or for other image related applications.</description>
<link>https://academictorrents.com/download/14cf27fca7f26714d2a5193dc95348a4712cdcdf</link>
</item>
<item>
<title>Twitch Emotes Images Dataset (Dataset)</title>
<description>This is a dataset containing over 1,200,000 images of twitch real twitch emotes. Most emotes (99.99%) are 28 by 28 Could be used to create a GAN or for other applications. Examples: </description>
<link>https://academictorrents.com/download/168649d9e29662e033d8db9c7bf0077c793d36c8</link>
</item>
<item>
<title>P. vivax (malaria) infected human blood smears (BBBC041) (Dataset)</title>
<description>### Description of the biological application Malaria is a disease caused by Plasmodium parasites that remains a major threat in global health, affecting 200 million people and causing 400,000 deaths a year. The main species of malaria that affect humans are Plasmodium falciparum and Plasmodium vivax. For malaria as well as other microbial infections, manual inspection of thick and thin blood smears by trained microscopists remains the gold standard for parasite detection and stage determination because of its low reagent and instrument cost and high flexibility. Despite manual inspection being extremely low throughput and susceptible to human bias, automatic counting software remains largely unused because of the wide range of variations in brightfield microscopy images. However, a robust automatic counting and cell classification solution would provide enormous benefits due to faster and more accurate quantitative results without human variability; researchers and medical professionals could better characterize stage-specific drug targets and better quantify patient reactions to drugs. Previous attempts to automate the process of identifying and quantifying malaria have not gained major traction partly due to difficulty of replication, comparison, and extension. Authors also rarely make their image sets available, which precludes replication of results and assessment of potential improvements. The lack of a standard set of images nor standard set of metrics used to report results has impeded the field. ### Images Images are in .png or .jpg format. There are 3 sets of images consisting of 1364 images (~80,000 cells) with different researchers having prepared each one: from Brazil (Stefanie Lopes), from Southeast Asia (Benoit Malleret), and time course (Gabriel Rangel). Blood smears were stained with Giemsa reagent. ### Ground truth The data consists of two classes of uninfected cells (RBCs and leukocytes) and four classes of infected cells (gametocytes, rings, trophozoites, and schizonts). Annotators were permitted to mark some cells as difficult if not clearly in one of the cell classes. The data had a heavy imbalance towards uninfected RBCs versus uninfected leukocytes and infected cells, making up over 95% of all cells. A class label and set of bounding box coordinates were given for each cell. For all data sets, infected cells were given a class label by Stefanie Lopes, malaria researcher at the Dr. Heitor Vieira Dourado Tropical Medicine Foundation hospital, indicating stage of development or marked as difficult. ### For more information These images were contributed by Jane Hung of MIT and the Broad Institute in Cambridge, MA. </description>
<link>https://academictorrents.com/download/2fed90eeaa0fbf98aba474c5d7e56f6290121507</link>
</item>
<item>
<title>DRIMDB (Diabetic Retinopathy Images Database) Database for Quality Testing of Retinal Images (Dataset)</title>
<description>Retinal image quality assessment (IQA) is a crucial process for automated retinal image analysis systems to obtain an accurate and successful diagnosis of retinal diseases. Consequently, the first step in a good retinal image analysis system is measuring the quality of the input image. We present an approach for finding medically suitable retinal images for retinal diagnosis. We used a three-class grading system that consists of good, bad, and outlier classes. We created a retinal image quality dataset with a total of 216 consecutive images called the Diabetic Retinopathy Image Database. We identified the suitable images within the good images for automatic retinal image analysis systems using a novel method. Subsequently, we evaluated our retinal image suitability approach using the Digital Retinal Images for Vessel Extraction and Standard Diabetic Retinopathy Database Calibration level 1 public datasets. The results were measured through the F1 metric, which is a harmonic mean of precision and recall metrics. The highest F1 scores of the IQA tests were 99.60%, 96.50%, and 85.00% for good, bad, and outlier classes, respectively. Additionally, the accuracy of our suitable image detection approach was 98.08%. Our approach can be integrated into any automatic retinal analysis system with sufficient performance scores. Good:  Bad:  Outlier: </description>
<link>https://academictorrents.com/download/99811ba62918f8e73791d21be29dcc372d660305</link>
</item>
<item>
<title>MS-Celeb-1M: {A} Dataset and Benchmark for Large-Scale Face Recognition (Dataset)</title>
<description>In this paper, we design a benchmark task and provide the associated datasets for recognizing face images and link them to corresponding entity keys in a knowledge base. More specifically, we propose a benchmark task to recognize one million celebrities from their face images, by using all the possibly collected face images of this individual on the web as training data. The rich information provided by the knowledge base helps to conduct disambiguation and improve the recognition accuracy, and contributes to various real-world applications, such as image captioning and news video analysis. Associated with this task, we design and provide concrete measurement set, evaluation protocol, as well as training data. We also present in details our experiment setup and report promising baseline results. Our benchmark task could lead to one of the largest classification problems in computer vision. To the best of our knowledge, our training dataset, which contains 10M images in version 1, is the largest publicly available one in the world.</description>
<link>https://academictorrents.com/download/9e67eb7cc23c9417f39778a8e06cca5e26196a97</link>
</item>
<item>
<title>Stanford Drone Dataset (Dataset)</title>
<description>When humans navigate a crowed space such as a university campus or the sidewalks of a busy street, they follow common sense rules based on social etiquette. In order to enable the design of new algorithms that can fully take advantage of these rules to better solve tasks such as target tracking or trajectory forecasting, we need to have access to better data. To that end, we contribute the very first large scale dataset (to the best of our knowledge) that collects images and videos of various types of agents (not just pedestrians, but also bicyclists, skateboarders, cars, buses, and golf carts) that navigate in a real world outdoor environment such as a university campus. In the above images, pedestrians are labeled in pink, bicyclists in red, skateboarders in orange, and cars in green.     ### CITATION If you find this dataset useful, please cite this paper (and refer the data as Stanford Drone Dataset or SDD): A. Robicquet, A. Sadeghian, A. Alahi, S. Savarese, Learning Social Etiquette: Human Trajectory Prediction In Crowded Scenes in European Conference on Computer Vision (ECCV), 2016.</description>
<link>https://academictorrents.com/download/01f95ea32e160e6c251ea55a87bd5a24b23cb03d</link>
</item>
<item>
<title>Inria Aerial Image Labeling Dataset (Dataset)</title>
<description>The Inria Aerial Image Labeling addresses a core topic in remote sensing: the automatic pixelwise labeling of aerial imagery. Dataset features: Coverage of 810 km² (405 km² for training and 405 km² for testing) Aerial orthorectified color imagery with a spatial resolution of 0.3 m Ground truth data for two semantic classes: building and not building (publicly disclosed only for the training subset) The images cover dissimilar urban settlements, ranging from densely populated areas (e.g., San Francisco’s financial district) to alpine towns (e.g,. Lienz in Austrian Tyrol). Instead of splitting adjacent portions of the same images into the training and test subsets, different cities are included in each of the subsets. For example, images over Chicago are included in the training set (and not on the test set) and images over San Francisco are included on the test set (and not on the training set). The ultimate goal of this dataset is to assess the generalization power of the techniques: while Chicago imagery may be used for training, the system should label aerial images over other regions, with varying illumination conditions, urban landscape and time of the year. The dataset was constructed by combining public domain imagery and public domain official building footprints.  Citation Emmanuel Maggiori, Yuliya Tarabalka, Guillaume Charpiat and Pierre Alliez. “Can Semantic Labeling Methods Generalize to Any City? The Inria Aerial Image Labeling Benchmark”. IEEE International Geoscience and Remote Sensing Symposium (IGARSS). 2017.</description>
<link>https://academictorrents.com/download/cf445f6073540af0803ee345f46294f088e7bba5</link>
</item>
<item>
<title>UCF Google Street View Dataset 2014 (Dataset)</title>
<description> The dataset contains 62,058 high quality Google Street View images. The images cover the downtown and neighboring areas of Pittsburgh, PA; Orlando, FL and partially Manhattan, NY. Accurate GPS coordinates of the images and their compass direction are provided as well. For each Street View placemark (i.e. each spot on one street), the 360° spherical view is broken down into 4 side views and 1 upward view. There is one additional image per placemark which shows some overlaid markers, such as the address, name of streets, etc. ### Citation: Please cite the following paper for which this data was collected (partially): Image Geo-localization based on Multiple Nearest Neighbor Feature Matching using Generalized Graphs. Amir Roshan Zamir and Mubarak Shah. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2014.</description>
<link>https://academictorrents.com/download/e52a8978af7c2f734f2b30795075dbcd50efc983</link>
</item>
<item>
<title>Flickr8k Dataset (Dataset)</title>
<description>8,000 photos and up to 5 captions for each photo. We introduce a new benchmark collection for sentence-based image description and search, consisting of 8,000 images that are each paired with five different captions which provide clear descriptions of the salient entities and events. … The images were chosen from six different Flickr groups, and tend not to contain any well-known people or locations, but were manually selected to depict a variety of scenes and situations  ## Citation Hodosh, Micah, Peter Young, and Julia Hockenmaier. "Framing image description as a ranking task: Data, models and evaluation metrics." Journal of Artificial Intelligence Research 47 (2013): 853-899.</description>
<link>https://academictorrents.com/download/9dea07ba660a722ae1008c4c8afdd303b6f6e53b</link>
</item>
<item>
<title>IDRiD (Indian Diabetic Retinopathy Image Dataset) (Dataset)</title>
<description>IDRiD (Indian Diabetic Retinopathy Image Dataset), is the first database representative of an Indian population. Moreover, it is the only dataset constituting typical diabetic retinopathy lesions and also normal retinal structures annotated at a pixel level. This dataset provides information on the disease severity of diabetic retinopathy, and diabetic macular edema for each image. This makes it perfect for development and evaluation of image analysis algorithms for early detection of diabetic retinopathy. This dataset was available as a part of "Diabetic Retinopathy: Segmentation and Grading Challenge" organised in conjuction with IEEE International Symposium on Biomedical Imaging (ISBI-2018), Washington D.C. The dataset is divided into three parts: A. Segmentation: It consists of 1. Original color fundus images (81 images divided into train and test set - JPG Files) 2. Groundtruth images for the Lesions (Microaneurysms, Haemorrhages, Hard Exudates and Soft Exudates divided into train and test set - TIF Files) and Optic Disc (divided into train and test set - TIF Files) B. Disease Grading: it consists of 1. Original color fundus images (516 images divided into train set (413 images) and test set (103 images) - JPG Files) 2. Groundtruth Labels for Diabetic Retinopathy and Diabetic Macular Edema Severity Grade (Divided into train and test set - CSV File) C. Localization: It consists of 1. Original color fundus images (516 images divided into train set (413 images) and test set (103 images) - JPG Files) 2. Groundtruth Labels for Optic Disc Center Location (Divided into train and test set - CSV File) 3. Groundtruth Labels for Fovea Center Location (Divided into train and test set - CSV File) For more information visit idrid.grand-challenge.org Sample images (scaled down)  Sample segmentations of microaneurysms (scaled down)  Paper:</description>
<link>https://academictorrents.com/download/3bb974ffdad31f9df9d26a63ed2aea2f1d789405</link>
</item>
<item>
<title>openem_example_data (Dataset)</title>
<description>Dataset to run OpenEM examples and tutorial.</description>
<link>https://academictorrents.com/download/b2a418e07b033bbb37ff46d030d9633d365c148e</link>
</item>
<item>
<title>The PatchCamelyon benchmark dataset (PCAM) (Dataset)</title>
<description>The PatchCamelyon benchmark is a new and challenging image classification dataset. It consists of 327.680 color images (96 x 96px) extracted from histopathologic scans of lymph node sections. Each image is annoted with a binary label indicating presence of metastatic tissue. PCam provides a new benchmark for machine learning models: bigger than CIFAR10, smaller than imagenet, trainable on a single GPU. ## Why PCam Fundamental machine learning advancements are predominantly evaluated on straight-forward natural-image classification datasets. Think MNIST, CIFAR, SVHN. Medical imaging is becoming one of the major applications of ML and we believe it deserves a spot on the list of go-to ML datasets. Both to challenge future work, and to steer developments into directions that are beneficial for this domain. We think PCam can play a role in this. It packs the clinically-relevant task of metastasis detection into a straight-forward binary image classification task, akin to CIFAR-10 and MNIST. Models can easily be trained on a single GPU in a couple hours, and achieve competitive scores in the Camelyon16 tasks of tumor detection and WSI diagnosis. Furthermore, the balance between task-difficulty and tractability makes it a prime suspect for fundamental machine learning research on topics as active learning, model uncertainty and explainability. </description>
<link>https://academictorrents.com/download/1561a180b11d4b746273b5ce46772ad36f1229b6</link>
</item>
<item>
<title>POLEN23E: image dataset for the Brazilian Savannah pollen types (Dataset)</title>
<description>The classification of pollen species and types is an important task in many areas like forensic palynology, archaeological palynology and melissopalynology. This paper presents the first annotated image dataset for the Brazilian Savannah pollen types that can be used to train and test computer vision based automatic pollen classifiers. A first baseline human and computer performance for this dataset has been established using 805 pollen images of 23 pollen types. In order to access the computer performance, a combination of three feature extractors and four machine learning techniques has been implemented, fine tuned and tested.  Citation: Gonçalves AB, Souza JS, Silva GGd, Cereda MP, Pott A, Naka MH, et al. (2016) Feature Extraction and Machine Learning for the Classification of Brazilian Savannah Pollen Grains. PLoS ONE 11(6): e0157044.  The link for the dataset is: .</description>
<link>https://academictorrents.com/download/ee51ec7708b35b023caba4230c871ae1fa254ab3</link>
</item>
<item>
<title>BRATS2013 Tumor-NoTumor Dataset (T-NT) (Dataset)</title>
<description>This dataset (called T-NT) contains images which contain or do not contain a tumor along with a segmentation of brain matter and the tumor. The goal is that it can be used to simulate bias in data in a controlled fashion. # Dataset Construction The synthetic data of the BRATS2013 dataset is used to construct this dataset. Each brain contains a tumor but it is typically only on one side. Only the right side is taken in order to have examples that do not have tumors. Each image is filtered to ensure it has enough brain in the image (more than 30% of the pixels). If the tumor takes up at least 1% of the pixels in the brain then it is considered to have a tumor. Here is an snippet from the code used to construct the dataset:     def get_labels(rightside):</description>
<link>https://academictorrents.com/download/d52ccc21455c7a82fd6e58964c89b7da99e0edf7</link>
</item>
<item>
<title>ImageClef - IAPR TC-12 Benchmark (Dataset)</title>
<description>The following archive contains the complete IAPR TC-12 Benchmark, which is now available free of charge and without any copyright restrictions. This is the most updated version of the IAPR TC-12 Benchmark and should be used from researchers from now on. This archive thereby comprises: 20000 images 1000 additional images previously used in object annotation tasks and/or the MUSCLE live event all complete (full-text) annotations (English, German, Random) all light annotations (English, German, Spanish, Random), i.e. all annotation tags except for the description tag Note: The image collection of the IAPR TC-12 Benchmark consists of 20,000 still natural images taken from locations around the world and comprising an assorted cross-section of still natural images. This includes pictures of different sports and actions, photographs of people, animals, cities, landscapes and many other aspects of contemporary life. Example images can be found in Section 2. Each image is associated with a text caption in up to three different languages (English, German and Spanish) . These annotations are stored in a database which is managed by a benchmark administration system that allows the specification of parameters according to which different subsets of the image collection can be generated. Section 3 provides more information and an annotation example. The IAPR TC-12 Benchmark is now available free of charge and without copyright restrictions. Information on how to access (and download) the complete benchmark as well as the resources used at ImageCLEFphoto 2006 - 2008 is given in Sections 4 and 5, while Section 6 provides links to related publications. 2 Collection Content The 20,000 images are high quality, multi-object, colour photographs that have been chosen according to strict image selection rules (see [2] for more details). Here are a couple of example images of some chosen categories: In publications based on the IAPR TC-12 Benchmark and/or the use of its data or a subset thereof, please cite the following publication: The IAPR Benchmark: A New Evaluation Resource for Visual Information Systems, Grubinger, Michael, Clough Paul D., Müller Henning, and Deselaers Thomas , International Conference on Language Resources and Evaluation, 24/05/2006, Genoa, Italy, (2006) Additional information on this data is available from the PhD thesis of Michael Grubinger: Michael Grubinger. Analysis and Evaluation of Visual Information Systems Performance. PhD Thesis. School of Computer Science and Mathematics, Faculty of Health, Engineering and Science, Victoria University, Melbourne, Australia, 2007. The thesis is available here:   Data can also be downloaded here:  ![]()</description>
<link>https://academictorrents.com/download/cf870b196222cf961a01c13999be9e4b7760cef1</link>
</item>
<item>
<title>COCO 2017 (Dataset)</title>
<description>Probably the most widely used dataset today for object localization is COCO: Common Objects in Context. Provided here are all the files from the 2017 version, along with an additional subset dataset created by fast.ai. Details of each COCO dataset is available from the COCO dataset page. The fast.ai subset contains all images that contain one of five selected categories, restricting objects to just those five categories; the categories are: chair couch tv remote book vase.</description>
<link>https://academictorrents.com/download/74dec1dd21ae4994dfd9069f9cb0443eb960c962</link>
</item>
<item>
<title>PASCAL Visual Object Classes (VOC) (Dataset)</title>
<description>Standardised image data sets for object class recognition - both 2007 and 2012 versions are provided here. The 2012 version has 20 classes. The train/val data has 11,530 images containing 27,450 ROI annotated objects and 6,929 segmentations.</description>
<link>https://academictorrents.com/download/e6d591cef9ea2840f7d8dfb6bb0e0503d5592128</link>
</item>
<item>
<title>Camvid: Motion-based Segmentation and Recognition Dataset (Dataset)</title>
<description>Segmentation dataset with per-pixel semantic segmentation of over 700 images, each inspected and confirmed by a second person for accuracy.</description>
<link>https://academictorrents.com/download/890e6716827f31cbd096c5aee7f777e30df7094a</link>
</item>
<item>
<title>Caltech 101 (Dataset)</title>
<description>A 37 category pet dataset with roughly 200 images for each class. The images have a large variations in scale, pose and lighting. Can also be used for localization.</description>
<link>https://academictorrents.com/download/85572f9564eb41b9d2045cbd35d742b2e3c4949f</link>
</item>
<item>
<title>Caltech-UCSD Birds-200-2011 (Dataset)</title>
<description>An image dataset with photos of 200 bird species (mostly North American); it can also be used for localization. Number of categories: 200; Number of images: 11,788; Annotations per image: 15 Part Locations, 312 Binary Attributes, 1 Bounding Box</description>
<link>https://academictorrents.com/download/1a5a206de443085ff07ca9a47530f1bfadf526ba</link>
</item>
<item>
<title>Oxford 102 Flowers (Dataset)</title>
<description>A 102 category dataset consisting of 102 flower categories, commonly occuring in the United Kingdom. Each class consists of 40 to 258 images. The images have large scale, pose and light variations.</description>
<link>https://academictorrents.com/download/9e3425fd8ba169e7ad499489152f06a6d5bb0be0</link>
</item>
<item>
<title>Stanford Cars (Dataset)</title>
<description>16,185 images of 196 classes of cars. The data is split into 8,144 training images and 8,041 testing images, where each class has been split roughly in a 50-50 split. Classes are typically at the level of Make, Model, Year</description>
<link>https://academictorrents.com/download/9c90b7f6208d430bff288845d45667ab2670da56</link>
</item>
<item>
<title>Oxford-IIIT Pet (Dataset)</title>
<description>A 37 category pet dataset with roughly 200 images for each class. The images have a large variations in scale, pose and lighting. Can also be used for localization.</description>
<link>https://academictorrents.com/download/97d27cad2802552b20b0c9a29dbf246bca73608d</link>
</item>
<item>
<title>MNIST (Dataset)</title>
<description>Classic dataset of small (28x28) handwritten grayscale digits, developed in the 1990s for testing the most sophisticated models of the day; today, often used as a basic “hello world” for introducing deep learning. This fast.ai datasets version uses a standard PNG format instead of the special binary format of the original, so you can use the regular data pipelines in most libraries; if you want to use just a single input channel like the original, simply pick a single slice from the channels axis</description>
<link>https://academictorrents.com/download/4d563087fb327739d7ec9ee9a0d32c4cb8b0355e</link>
</item>
<item>
<title>CIFAR100 (Dataset)</title>
<description>This dataset is just like the CIFAR-10, except it has 100 classes containing 600 images each. There are 500 training images and 100 testing images per class. The 100 classes in the CIFAR-100 are grouped into 20 superclasses. Each image comes with a “fine” label (the class to which it belongs) and a “coarse” label (the superclass to which it belongs).</description>
<link>https://academictorrents.com/download/4fb115df73d3313fae9264fd6c0bad061add2d63</link>
</item>
<item>
<title>CIFAR10 (Dataset)</title>
<description>60000 32x32 colour images in 10 classes, with 6000 images per class (50000 training images and 10000 test images). Very widely used today for testing performance of new algorithms. This fast.ai datasets version uses a standard PNG format instead of the platform-specific binary formats of the original, so you can use the regular data pipelines in most libraries</description>
<link>https://academictorrents.com/download/e5834a981a7337f81fe5bad6c890c949c73ca30c</link>
</item>
<item>
<title>nyu_depth_v2_labeled.mat (Dataset)</title>
<description>The labeled dataset is a subset of the Raw Dataset. It is comprised of pairs of RGB and Depth frames that have been synchronized and annotated with dense labels for every image. In addition to the projected depth maps, we have included a set of preprocessed depth maps whose missing values have been filled in using the colorization scheme of Levin et al. Unlike, the Raw dataset, the labeled dataset is provided as a Matlab .mat file with the following variables: accelData – Nx4 matrix of accelerometer values indicated when each frame was taken. The columns contain the roll, yaw, pitch and tilt angle of the device. depths – HxWxN matrix of in-painted depth maps where H and W are the height and width, respectively and N is the number of images. The values of the depth elements are in meters. images – HxWx3xN matrix of RGB images where H and W are the height and width, respectively, and N is the number of images. instances – HxWxN matrix of instance maps. Use get_instance_masks.m in the Toolbox to recover masks for each object instance in a scene. labels – HxWxN matrix of object label masks where H and W are the height and width, respectively and N is the number of images. The labels range from 1..C where C is the total number of classes. If a pixel’s label value is 0, then that pixel is ‘unlabeled’. names – Cx1 cell array of the english names of each class. namesToIds – map from english label names to class IDs (with C key-value pairs) rawDepths – HxWxN matrix of raw depth maps where H and W are the height and width, respectively, and N is the number of images. These depth maps capture the depth images after they have been projected onto the RGB image plane but before the missing depth values have been filled in. Additionally, the depth non-linearity from the Kinect device has been removed and the values of each depth image are in meters. rawDepthFilenames – Nx1 cell array of the filenames (in the Raw dataset) that were used for each of the depth images in the labeled dataset. rawRgbFilenames – Nx1 cell array of the filenames (in the Raw dataset) that were used for each of the RGB images in the labeled dataset. scenes – Nx1 cell array of the name of the scene from which each image was taken. sceneTypes – Nx1 cell array of the scene type from which each image was taken.</description>
<link>https://academictorrents.com/download/47a9a46bb784b394e398228d4c85a8d61d01dfa8</link>
</item>
<item>
<title>CamVid - The Cambridge-driving Labeled Video Database (Dataset)</title>
<description>The Cambridge-driving Labeled Video Database (CamVid) is the first collection of videos with object class semantic labels, complete with metadata. The database provides ground truth labels that associate each pixel with one of 32 semantic classes. The database addresses the need for experimental data to quantitatively evaluate emerging algorithms. While most videos are filmed with fixed-position CCTV-style cameras, our data was captured from the perspective of a driving automobile. The driving scenario increases the number and heterogeneity of the observed object classes. Over ten minutes of high quality 30Hz footage is being provided, with corresponding semantically labeled images at 1Hz and in part, 15Hz. The CamVid Database offers four contributions that are relevant to object analysis researchers. First, the per-pixel semantic segmentation of over 700 images was specified manually, and was then inspected and confirmed by a second person for accuracy. Second, the high-quality and large resolution color video images in the database represent valuable extended duration digitized footage to those interested in driving scenarios or ego-motion. Third, we filmed calibration sequences for the camera color response and intrinsics, and computed a 3D camera pose for each frame in the sequences. Finally, in support of expanding this or other databases, we offer custom-made labeling software for assisting users who wish to paint precise class-labels for other images and videos. We evaluated the relevance of the database by measuring the performance of an algorithm from each of three distinct domains: multi-class object recognition, pedestrian detection, and label propagation. ![]() #### Citation Request Segmentation and Recognition Using Structure from Motion Point Clouds, ECCV 2008 Brostow, Shotton, Fauqueur, Cipolla Semantic Object Classes in Video: A High-Definition Ground Truth Database Pattern Recognition Letters Brostow, Fauqueur, Cipolla</description>
<link>https://academictorrents.com/download/a6431e7dd33615194c5936fd8a35db043ab51058</link>
</item>
<item>
<title>Holistic Recognition of Low Quality License Plates (HDR dataset) (Dataset)</title>
<description>This dataset focuses on recognition of license plates in low resolution and low quality images. ![]() ### Citation Request J. Špaňhel, J. Sochor, R. Juránek, A. Herout, L. Maršík and P. Zemčík, "Holistic recognition of low quality license plates by CNN using track annotated data," 2017 14th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Lecce, 2017, pp. 1-6. doi: 10.1109/AVSS.2017.8078501</description>
<link>https://academictorrents.com/download/8ed33d02d6b36c389dd077ea2478cc83ad117ef3</link>
</item>
<item>
<title>VizWiz v1.0 dataset (Answering Visual Questions from Blind People) (Dataset)</title>
<description>We propose an artificial intelligence challenge to design algorithms that assist people who are blind to overcome their daily visual challenges. For this purpose, we introduce the VizWiz dataset, which originates from a natural visual question answering setting where blind people each took an image and recorded a spoken question about it, together with 10 crowdsourced answers per visual question. Our proposed challenge addresses the following two tasks for this dataset: (1) predict the answer to a visual question and (2) predict whether a visual question cannot be answered. Ultimately, we hope this work will educate more people about the technological needs of blind people while providing an exciting new opportunity for researchers to develop assistive technologies that eliminate accessibility barriers for blind people.     VizWiz v1.0 dataset download: 20,000 training image/question pairs 200,000 training answer/answer confidence pairs 3,173 image/question pairs 31,730 validation answer/answer confidence pairs 8,000 image/question pairs Python API to read and visualize the VizWiz dataset Python challenge evaluation code     ![]() ### Publications Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. "VizWiz Grand Challenge: Answering Visual Questions from Blind People." IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. Jeffrey P. Bigham, Chandrika Jayant, Hanjie Ji, Greg Little, Andrew Miller, Robert C. Miller, Robin Miller, Aubrey Tatarowicz, Brandyn White, Samuel White, and Tom Yeh. "VizWiz: Nearly Real-time Answers to Visual Questions." ACM User Interface Software and Technology Symposium (UIST), 2010.</description>
<link>https://academictorrents.com/download/b633e14aa084fab57f20ad0b4612e0932ae1f2dc</link>
</item>
<item>
<title>Human MCF7 cells – compound-profiling experiment (BBBC021v1) (Dataset)</title>
<description>![]() ![]() ### Description of the biological application Phenotypic profiling attempts to summarize multiparametric, feature-based analysis of cellular phenotypes of each sample so that similarities between profiles reflect similarities between samples. Profiling is well established for biological readouts such as transcript expression and proteomics. Image-based profiling, however, is still an emerging technology. This image set provides a basis for testing image-based profiling methods wrt. to their ability to predict the mechanisms of action of a compendium of drugs. The image set was collected using a typical set of morphological labels and uses a physiologically relevant p53-wildtype breast-cancer model system (MCF-7) and a mechanistically distinct set of targeted and cancer-relevant cytotoxic compounds that induces a broad range of gross and subtle phenotypes. ### Images The images are of MCF-7 breast cancer cells treated for 24 h with a collection of 113 small molecules at eight concentrations. The cells were fixed, labeled for DNA, F-actin, and Β-tubulin, and imaged by fluorescent microscopy as described [Caie et al. Molecular Cancer Therapeutics, 2010]. There are 39,600 image files (13,200 fields of view imaged in three channels) in TIFF format. We provide the images in 55 ZIP archives, one for each microtiter plate. ### Metadata The file BBBC021_v1_image.csv contains the metadata, with the following fields:     TableNumber ImageNumber Image_FileName_DAPI Image_PathName_DAPI Image_FileName_Tubulin Image_PathName_Tubulin Image_FileName_Actin Image_PathName_Actin Image_Metadata_Plate_DAPI Image_Metadata_Well_DAPI Replicate Image_Metadata_Compound Image_Metadata_Concentration     ### Ground truth B A subset of the compound-concentrations have been identified as clearly having one of 12 different primary mechanims of action. mechanistic classes were selected so as to represent a wide cross-section of cellular morphological phenotypes. The differences between phenotypes were in some cases very subtle: we were only able to identify 6 of the 12 mechanisms visually; the remainder were defined based on the literature. The file BBBC021_v1_moa.csv contains the mechanisms of action of 103 compound-concentrations (38 compounds at 1–7 concentrations each). The fields are:     compound concentration moa     ### Recommended citation "We used image set BBBC021v1 [Caie et al., Molecular Cancer Therapeutics, 2010], available from the Broad Bioimage Benchmark Collection [Ljosa et al., Nature Methods, 2012]."</description>
<link>https://academictorrents.com/download/014980e8a505760ed4c33641ac7e603d6e1778f4</link>
</item>
<item>
<title>MICCAI 2013 Challenge on Multimodal Brain Tumor Segmentation (BraTS2013)	 (Dataset)</title>
<description>A publicly available set of training data can be downloaded for algorithmic tweaking and tuning from the Virtual Skeleton Database. The training data consists of multi-contrast MR scans of 30 glioma patients (both low-grade and high-grade, and both with and without resection) along with expert annotations for "active tumor" and "edema". For each patient, T1, T2, FLAIR, and post-Gadolinium T1 MR images are available. All volumes were linearly co-registered to the T1 contrast image, skull stripped, and interpolated to 1mm isotropic resolution. No attempt was made to put the individual patients in a common reference space. The MR scans, as well as the corresponding reference segmentations, are distributed in the ITK- and VTK-compatible MetaIO file format. Patients with high- and low-grade gliomas have file names "BRATS_HG" and "BRATS_LG", respectively. All images are stored as signed 16-bit integers, but only positive values are used. The manual segmentations (file names ending in "_truth.mha") have only five intensity levels: 1 for Non-brain, non-tumor, necrosis, cyst, hemorrhage, 2 for Surrounding edema, 3 for Non-enhancing tumor, 4 for enhancing tumor core and 0 for everything else. Detailed technical documentation on the used MetaIO file format is available here. The training data also contains simulated images for 25 high-grade and 25 low-grade glioma subjects. These simulated images closely follow the conventions used for the real data, except that their file names start with "SimBRATS"; they are all in BrainWeb space; and their MR scans and ground truth segmentations are stored using unsigned 16 bit and unsigned 8 bit integers, respectively. Details on the simulation method are available here. Testing data A set of independent testing data will be provided on the day of the challenge itself. This testing data will be similar to the training data, except that the reference segmentation will not be made publicly available. ![]()</description>
<link>https://academictorrents.com/download/39c5a52bda7b5b701cecfc454a79d385868d4f3d</link>
</item>
<item>
<title>Caudate Segmentation Evaluation 2007 (CAUSE07) (Dataset)</title>
<description>CAUSE07 is a competition that was held as part of the workshop 3D Segmentation in the Clinic: A Grand Challenge, on October 26, 2007 in conjunction with MICCAI 2007. The goal of this competition was to compare different algorithms to segment the caudate nucleaus from brain MRI scans. Through this website, the competition continues. You can browse the results of various systems, and read papers and descriptions about the methods that have been applied to the CAUSE07 data set. If you want to join the competition, you can register a team, download training and test data, and submit the results of your own algorithms, provided you adhere to and agree with the rules. More information is available in the answers to frequently asked questions. ![]() ## Citation Request "3D Segmentation in the Clinic: A Grand Challenge", B. van Ginneken, T. Heimann, and M. Styner. In: T. Heimann, M. Styner, B. van Ginneken (Eds.): 3D Segmentation in the Clinic: A Grand Challenge, pp. 7-15, 2007.</description>
<link>https://academictorrents.com/download/d6c066ef308cc704c8898d5f87cf55e986475fb5</link>
</item>
<item>
<title>Breast Cancer Cell Segmentation (Dataset)</title>
<description>There are about 58 H&amp;E stained histopathology images used in breast cancer cell detection with associated ground truth data available. Routine histology uses the stain combination of hematoxylin and eosin, commonly referred to as H&amp;E. These images are stained since most cells are essentially transparent, with little or no intrinsic pigment. Certain special stains, which bind selectively to particular components, are be used to identify biological structures such as cells. In those images, the challenging problem is cell segmentation for subsequent classification in benign and malignant cells. The ground truth have been obtained for one image containing benign cells. | Image: |Ground Truth: | |&amp;mdash;-|&amp;mdash;-| | ![]() | ![]() | All images: ![]()</description>
<link>https://academictorrents.com/download/b79869ca12787166de88311ca1f28e3ebec12dec</link>
</item>
<item>
<title>Animals with Attributes 2 (AwA2) dataset (Dataset)</title>
<description>This dataset provides a platform to benchmark transfer-learning algorithms, in particular attribute base classification and zero-shot learning [1]. It can act as a drop-in replacement to the original Animals with Attributes (AwA) dataset [2,3], as it has the same class structure and almost the same characteristics. It consists of 37322 images of 50 animals classes with pre-extracted feature representations for each image. The classes are aligned with Osherson s classical class/attribute matrix [3,4], thereby providing 85 numeric attribute values for each class. Using the shared attributes, it is possible to transfer information between different classes. The image data was collected from public sources, such as Flickr, in 2016. In the process we made sure to only include images that are licensed for free use and redistribution, please see the archive for the individual license files. ![]() ### Publications Please cite the following paper when using the dataset: [1] Y. Xian, C. H. Lampert, B. Schiele, Z. Akata. "Zero-Shot Learning - A Comprehensive Evaluation of the Good, the Bad and the Ugly" arXiv:1707.00600 [cs.CV] Attribute based classification and the original Animals with Attributes (AwA) data is described in: [2] C. H. Lampert, H. Nickisch, and S. Harmeling. "Learning To Detect Unseen Object Classes by Between-Class Attribute Transfer". In CVPR, 2009 (pdf) [3] C. H. Lampert, H. Nickisch, and S. Harmeling. "Attribute-Based Classification for Zero-Shot Visual Object Categorization". IEEE T-PAMI, 2013 (pdf) The class/attribute matrix was originally created by: [4] D. N. Osherson, J. Stern, O. Wilkie, M. Stob, and E. E. Smith. "Default probability". Cognitive Science, 15(2), 1991. [5] C. Kemp, J. B. Tenenbaum, T. L. Griffiths, T. Yamada, and N. Ueda. "Learning systems of concepts with an infinite relational model". In AAAI, 2006.</description>
<link>https://academictorrents.com/download/1490aec815141cdb50a32b81ef78b1eaf6b38b03</link>
</item>
<item>
<title>UrbanMapper 3D (Digital Surface Model and Digital Terrain Model) Dataset (Dataset)</title>
<description>Competitors will receive an orthorectified color image, Digital Surface Model (DSM), and Digital Terrain Model (DTM) for each geographic area of interest (AOI). The DSM indicates the height of the earth, with objects such as buildings and trees included. The DTM indicates only the height of the ground. Both should be expected to include some errors, and errors may be expected to be similar in the provisional and sequestered data sets. The difference in the DSM and DTM indicates height of objects above ground. All input files provided are raster GeoTIFF images. Ground truth building labels will also be provided for a subset of the data to be used for training ![]() ![]()</description>
<link>https://academictorrents.com/download/4ccd3743861d827ac80f0d2b234d7fcfdad2a31d</link>
</item>
<item>
<title>UC Merced Land Use Dataset (Dataset)</title>
<description>This is a 21 class land use image dataset meant for research purposes. There are 100 images for each of the following classes:     agricultural airplane baseballdiamond beach buildings chaparral denseresidential forest freeway golfcourse harbor intersection mediumresidential mobilehomepark overpass parkinglot river runway sparseresidential storagetanks tenniscourt     Each image measures 256x256 pixels. ![]() The images were manually extracted from large images from the USGS National Map Urban Area Imagery collection for various urban areas around the country. The pixel resolution of this public domain imagery is 1 foot. Please cite the following paper when publishing results that use this dataset: Yi Yang and Shawn Newsam, "Bag-Of-Visual-Words and Spatial Extensions for Land-Use Classification," ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems (ACM GIS), 2010. Shawn D. Newsam Assistant Professor and Founding Faculty Electrical Engineering &amp; Computer Science University of California, Merced Email: snewsam@ucmerced.edu Web:  This material is based upon work supported by the National Science Foundation under Grant No. 0917069.</description>
<link>https://academictorrents.com/download/e9ac5edf285a43309e57e1289e8816a4e78a937c</link>
</item>
<item>
<title>AVA: A Large-Scale Database for Aesthetic Visual Analysis (Dataset)</title>
<description>Aesthetic Visual Analysis (AVA) contains over 250,000 images along with a rich variety of meta-data including a large number of aesthetic scores for each image, semantic labels for over 60 categories as well as labels related to photographic style for high-level image quality categorization.</description>
<link>https://academictorrents.com/download/71631f83b11d3d79d8f84efe0a7e12f0ac001460</link>
</item>
<item>
<title>Small Object Dataset (Dataset)</title>
<description>Images of small objects for small instance detections.  Currently four object types are available. ![]() We collect four datasets of small objects from images/videos on the Internet (e.g.YouTube or Google). Fly Dataset: contains 600 video frames with an average of 86 ± 39 flies per frame (648×72 @ 30 fps). 32 images are used for training (1:6:187) and 50 images for testing (301:6:600). Honeybee Dataset: contains 118 images with an average of 28 ± 6 honeybees per image (640×480). The dataset is divided evenly for training and test sets. Only the first 32 images are used for training. Fish Dataset: contains 387 frames of video with an average of 56±9 fish per frame (300×410 @ 30 fps). 32 images are used for training (1:3:94) and 65 for testing (193:3:387). Seagull Dataset: contains three high-resolution images (624×964) with an average of 866±107 seagulls per image. The first image is used for training, and the rest for testing. Cite this paper: </description>
<link>https://academictorrents.com/download/8e751c111cf90123374b5f0cf61e6af9f5e5231e</link>
</item>
<item>
<title>Downsampled ImageNet 32x32 (Dataset)</title>
<description>This page includes downsampled ImageNet images, which can be used for density estimation and generative modeling experiments. Images come in two resolutions: 32x32 and 64x64, and were introduced in Pixel Recurrent Neural Networks. Please refer to the Pixel RNN paper for more details and results. ![]()</description>
<link>https://academictorrents.com/download/bf62f5051ef878b9c357e6221e879629a9b4b172</link>
</item>
<item>
<title>Downsampled ImageNet 64x64 (Dataset)</title>
<description>This page includes downsampled ImageNet images, which can be used for density estimation and generative modeling experiments. Images come in two resolutions: 32x32 and 64x64, and were introduced in Pixel Recurrent Neural Networks. Please refer to the Pixel RNN paper for more details and results. ![]()</description>
<link>https://academictorrents.com/download/96816a530ee002254d29bf7a61c0c158d3dedc3b</link>
</item>
<item>
<title>VGG Cell Dataset from Learning To Count Objects in Images   (Dataset)</title>
<description>![]() We generated a dataset of  200 images, and used random subsets of the first 100 images to perform training and parameter validations, and the second 100 images to test the counting accuracy. Below, we show some representative results for cell counting for the previously unseen images ### Acknowledgements This work is a part of the EU VisRec project (ERC grant VisRec no. 228180).</description>
<link>https://academictorrents.com/download/b32305598175bb8e03c5f350e962d772a910641c</link>
</item>
<item>
<title>NIST 8-Bit Gray Scale Images of Fingerprint Image Groups (FIGS) (Dataset)</title>
<description>The NIST database of fingerprint images contains 2000 8-bit gray scale fingerprint image pairs. Each image is 512-by-512 pixels with 32 rows of white space at the bottom and classified using one of the five following classes:</description>
<link>https://academictorrents.com/download/d7e67e86f0f936773f217dbbb9c149c4d98748c6</link>
</item>
<item>
<title>Movies Fight Detection Dataset (Dataset)</title>
<description>Whereas the action recognition community has focused mostly on detecting simple actions like clapping, walking or jogging, the detection of fights or in general aggressive behaviors has been comparatively less studied. Such capability may be extremely useful in some video surveillance scenarios like in prisons, psychiatric or elderly centers or even in camera phones. After an analysis of previous approaches we test the well-known Bag-of-Words framework used for action recognition in the specific problem of fight detection, along with two of the best action descriptors currently available: STIP and MoSIFT. For the purpose of evaluation and to foster research on violence detection in video we introduce a new video database containing 1000 sequences divided in two groups: fights and non-fights. Experiments on this database and another one with fights from action movies show that fights can be detected with near 90% accuracy.</description>
<link>https://academictorrents.com/download/70e0794e2292fc051a13f05ea6f5b6c16f3d3635</link>
</item>
<item>
<title>Labeled Fishes in the Wild (Dataset)</title>
<description>The labeled fishes in the wild image dataset is provided by NOAA Fisheries (National Marine Fisheries Service) to encourage development, testing, and performance assessment of automated image analysis algorithms for unconstrained underwater imagery. The dataset includes images of fish, invertebrates, and the seabed that were collected using camera systems deployed on a remotely operated vehicle (ROV) for fisheries surveys. Annotation data are included in accompanying data files (.dat, .vec, and .info) that describe the locations of the marked fish targets in the images. The manuscript (Cutter et al., 2015) demonstrates methods for automated detection of fish based on classifiers developed using the training image dataset, and evaluated using the test set. This dataset is offered for further development of detection of fish or invertebrates in complex environments; tracking of multiple animal targets in video image sequences; recognition and classification of animal species; measurement of animals in stereo image pairs; and characterization of seabed habitats. Recommended citation: Cutter, G.; Stierhoff, K.; Zeng, J. (2015) "Automated detection of rockfish in unconstrained underwater videos using Haar cascades and a new image dataset: labeled fishes in the wild," IEEE Winter Conference on Applications of Computer Vision Workshops, pp. 57-62. The NOAA scientists who are stewards of these data may have archives of images that can provide additional opportunities for collaboration to apply and assess algorithms. Credit for use of these datasets should be provided in publications, as described in the “how-to-cite.txt” documents included in the dataset archive or as shown above. ##Dataset Labeled Fishes in the Wild image dataset (v. 1.1) (Download 423 MB). Labeled fishes in the wild has three components: a training and validation positive image set (verified fish), a negative image set (non-fish), and a test image set. The training and test sets have accompanying annotation data that define the location and extent of each marked fish target object in the images. These represent bounding rectangles defined by expert analysts, and are in the format of .dat files used by OpenCV. Training and validation positive image set: contains images of rockfish (Sebastes spp.) and other associated species near the seabed, collected using a forward-oblique-looking digital still camera deployed on a remotely operated vehicle (ROV) by the Southwest Fisheries Science Center during surveys of rocky seabed environments offshore of southern California. Still frames from these cameras represent instances during a survey where the ROV was moving slowly, and motion effects are not a factor. The training set comprises 929 image files, containing 1005 marked fish with associated annotations (their marked locations and bounding rectangles). The marks define fish of various species, sizes, and ranges to the camera, and includes portions of different background composition. Training and validation negative image set: includes 3167 images. The 147 seabed negative images provided in the downloadable archive were extracted from the labeled fishes in the wild training and test image sets (regions containing no fish were extracted). The remaining 3020 images are available from the tutorial on OpenCV HaarTraining, and available from the data negatives directory. Test image set: contains an image sequence collected using the ROV’s high-definition (HD; 1080i) video camera during a near-seabed survey of fish. The test imagery for detection comprises video footage from ROV surveys. The video clip (“TEST_VIDEO_ROV10.mp4”; 210 frames at 3 frames per second (fps)) used to evaluate detectors for this study represents every 10th frame of the original video sequence (2-minute duration, approximately 30 fps). All fish targets are annotated for the 210-frame, 3fps test video. Annotations of fish in the test video include a descriptor, “verified” or “apparent,” where verified indicates that a video analyst could identify the fish as such, and apparent objects were believed to be fish, but were not verifiable based on attributes visible in a single frame. These apparent fish may appear as faint blobs in the distance. These distinctions are made in the annotation data because we believe that some classifiers will detect these apparent fish, but we do not expect the classifier to do so; nor do we necessarily want the detector to do so. That is, if a classifier is detecting those apparent fish, then it is probably detecting many other non-fish targets in the images, thereby making it inefficient and impractical. A total of 2061 fish objects were marked in the annotated frames of the dataset test video. Of those, 1008 were verified fish, and 1053 were apparent fish. During the sequence the ROV is moving; the background appears to be moving and is illuminated from different directions (as the ROV moves and rotates); small particles in the water current stream past; fish are still or moving at various speeds; fish are oriented in many directions; some fish are hidden partially behind rocks or in crevices; some indistinct fish-like objects appear in the distance. The original Labeled fishes in the wild dataset (v1.0, Dec. 2014) contained only the decimated test video sequence ("Test_ROV_video_h264_decim.mp4") that contained only the marked frames from the original video. One tenth of the frames of the full frame-rate video were marked for locations of fish targets. This version of the dataset (v1.1, Jan. 2015) also contains the full test video sequence ("Test_ROV_video_h264_full.mp4"). Both the full and decimated videos have accompanying text files with analyst marks (following OpenCV .dat file conventions). Generally, for m marks, the format is: Video-filename(frame#) #-of-marks x1 y1 w1 h1 x2 y2 w2 h2 ... xm ym wm hm. For example, in the case of two marks, the final eight values define the bounding rectangles: Test_ROV_video_h264_full.mp4(fr_14) 2 1021 362 94 63 953 289 90 61. The marks file for the decimated video ("Test_ROV_video_h264_decim_marks.dat") indicates the frame number for the decimated and full sequence, e.g. Test_ROV_video_h264_decim.mp4(fr_1)(fullfr_14) 2 1021 362 94 63 953 289 90 61. There are 2101 frames in the full video and 210 frames in the decimated video, but 206 frames were marked; i.e. a few of the examined frames did not contain fish. Contact: george dot cutter at noaa dot gov. ![]()</description>
<link>https://academictorrents.com/download/41bc10c77d54b49fb0a96ff5d4a0814bc2ab7da7</link>
</item>
<item>
<title>Tiny Images Dataset (Dataset)</title>
<description> ## Overview This page has links for downloading the Tiny Images dataset, which consists of 79,302,017 images, each being a 32x32 color image. This data is stored in the form of large binary files which can be accessed by a Matlab toolbox that we have written. You will need around 400Gb of free disk space to store all the files. In total there are 5 files that need to be downloaded, 3 of which are large binary files consisting of (i) the images themselves; (ii) their associated metadata (filename, search engine used, ranking etc.); (iii) Gist descriptors for each image. The other two files are the Matlab toolbox and index data file that together let you easily load in data from the binaries. ## Downloads Note that these files are very large and will take a considerable time to download. Please ensure you have sufficient disk space before commencing the download. 1. Image binary (227Gb) 2. Metadata binary (57Gb) 3. Gist binary (114Gb) 4. Index data (7Mb) 5. Matlab Tiny Images toolbox (150Kb) ## Overview The 79 million images are stored in one giant binary file, 227Gb in size. The metadata accompanying each image is also in a single giant file, 57Gb in size. To read images/metadata from these files, we have provided some Matlab wrapper functions. There are two versions of the functions for reading image data: * (i) loadTinyImages.m - plain Matlab function (no MEX), runs under 32/64bits. Loads images in by image number. Use this by default. * (ii) read_tiny_big_binary.m - Matlab wrapper for 64-bit MEX function. A bit faster and more flexible than (i), but requires a 64-bit machine. There are two types of annotation data: * (i) Manual annotation data, sorted in annotations.txt, that holds the label of images manually inspected to see if image content agrees with noun used to collect it. Some other information, such as search engine, is also stored. This data is available for only a very small portion of images. * (ii) Automatic annotation data, stored in tiny_metadata.bin, consisting of information relating the gathering of the image, e.g. search engine, which page, url to thumbnail etc. This data is available for all 79 million images. ## Requirements 1. Around 300Gb of disk space. 2. If you want to use the MEX versions of the code for reading in the data, you will need a 64-bit machine. But for most purposes, the Matlab implementation (loadTinyImages.m), which can use either 32 or 64bits will work perfectly well. To discover if you have a 32/64bit machine, type  uname -a  in an xterm (if using linux). ## Files The .tgz file should contain 10 files 1. loadTinyImages.m &amp;mdash; read tiny image data, pure Matlab version. 2. loadGroundTruth.m &amp;mdash; read annotations.txt file holding manual annotations 3. read_tiny_big_binary.m &amp;mdash; read tiny image data, 64-bit Matlab/MEX version 4. read_tiny_big_metadata.m &amp;mdash; read tiny image metadata, 64-bit Matlab/MEX version 5. read_tiny_gist_binary.m &amp;mdash; read tiny Gist, 64-bit Matlab/MEX version 6. read_tiny_binary_big_core.c &amp;mdash; 64-bit MEX source code for image reading 7. read_tiny_metadata_big_core.c &amp;mdash; 64-bit MEX source code for metadata reading 8. read_tiny_binary_gist_core.c &amp;mdash; 64-bit MEX source code for gist reading 9. compute_hash_function.m &amp;mdash; utility function to do fast string searching as used by read_tiny_big_binary.m and read_tiny_big_metadata.m 10. fast_str2num.m &amp;mdash; utility function for &amp;mdash; &amp;mdash; read_tiny_big_metadata.m 11. annotations.txt &amp;mdash; text file holding list of annotated images 12. README.txt &amp;mdash; this file Separately, you should have downloaded the following files 1. tiny_images.bin - 227Gb file holding 79,302,017 images 2. tiny_metadata.bin - 57Gb file holding metadata for all 79,302,017 images 3. tinygist80million.bin - 114Gb file holding 384-dim Gist descriptors for all 79,302,017 images 4. tiny_index.mat - 7Mb file holding index info, including: word - cell array of all 75,846 nouns for which we have images in tiny_images.bin num_imgs - vector with #images per noun for all 75,846 nouns ## Preliminaries Before the functions can be used you must do two things: 1. Set the absolute paths to the binary files in the Matlab functions. There are a total of 7 lines that must be set: *(i) loadTinyImages.m, line 14 &amp;mdash; set path to tiny_images.bin file *(ii) read_tiny_big_binary.m, line 40 &amp;mdash; set path to tiny_images.bin file *(iii) read_tiny_big_binary.m, line 42 &amp;mdash; set path to tiny_index.mat file *(iv) read_tiny_big_metadata.m, line 63 &amp;mdash; set path to tiny_metadata.bin file *(v) read_tiny_big_metadata.m, line 65 &amp;mdash; set path to tiny_index.mat file *(vi) read_tiny_gist_binary.m, line 36 &amp;mdash; set path to tiny_index.mat file *(vii) read_tiny_gist_binary.m, line 38 &amp;mdash; set path to tiny_metadata.bin file 2. If using the MEX versions, they must be compiled with the commands: *(i) mex read_tiny_binary_big_core.c *(ii) mex read_tiny_metadata_big_core.c *(iii) mex read_tiny_binary_gist_core.c ## Usage Here are some examples of the scripts in use. Please look at the comments at the top of each file for more extensive explanations. ## loadTinyImages.m</description>
<link>https://academictorrents.com/download/03b779ffefa8efc30c2153f3330bb495bdc3e034</link>
</item>
<item>
<title>Labeled Faces in the Wild aligned (LFW-a) (Dataset)</title>
<description>The "Labeled Faces in the Wild-a" image collection is a database of labeled, face images intended for studying Face Recognition in unconstrained images. It contains the same images available in the original Labeled Faces in the Wild data set, however, here we provide them after alignment using a commercial face alignment software. Some of our results, published in [1,2,3], were produced using these images. We show this alignment to improve the performance of face recognition algorithms. More information on how these images were aligned may be found in the two papers. We have maintained the same directory structure as in the original LFW data set, and so these images can be used as direct substitutes for those in the original image set. Note, however, that the images available here are grayscale versions of the originals. Citation: If you find these images useful and use them in your work, please follow these guidlines: Comply with any instructions specified for the original LFW data set Cite one (or all) of the papers [1,2,3] below References: [1] Lior Wolf, Tal Hassner, and Yaniv Taigman, Effective Face Recognition by Combining Multiple Descriptors and Learned Background Statistics, IEEE Trans. on Pattern Analysis and Machine Intelligence (TPAMI), 33(10), Oct. 2011 (PDF) [2] Lior Wolf, Tal Hassner and Yaniv Taigman, Similarity Scores based on Background Samples, Asian Conference on Computer Vision (ACCV), Xi  an, Sept 2009 (PDF) [3] Yaniv Taigman, Lior Wolf and Tal Hassner, Multiple One-Shots for Utilizing Class Label Information, The British Machine Vision Conference (BMVC), London, Sept 2009 (project, PDF)</description>
<link>https://academictorrents.com/download/403e6d6945a64dd1b9e185a6cd8d029274efccdc</link>
</item>
<item>
<title>Labeled Faces in the Wild - aligned with funneling (Dataset)</title>
<description>Data from the paper: "Unsupervised Joint Alignment of Complex Images" Gary B. Huang and Vidit Jain and Erik Learned-Miller ICCV 2007 Welcome to Labeled Faces in the Wild, a database of face photographs designed for studying the problem of unconstrained face recognition. The data set contains more than 13,000 images of faces collected from the web. Each face has been labeled with the name of the person pictured. 1680 of the people pictured have two or more distinct photos in the data set. The only constraint on these faces is that they were detected by the Viola-Jones face detector. More details can be found in the technical report below. Information: 13233 images 5749 people 1680 people with two or more images</description>
<link>https://academictorrents.com/download/073ecac13cf175ddc5617c1f2897b15ca7accd59</link>
</item>
<item>
<title>Labeled Faces in the Wild - aligned with deep funneling (Dataset)</title>
<description>Data from the paper: Learning to Align from Scratch Gary B. Huang and Marwan Mattar and Honglak Lee and Erik Learned-Miller NIPS 2012 Welcome to Labeled Faces in the Wild, a database of face photographs designed for studying the problem of unconstrained face recognition. The data set contains more than 13,000 images of faces collected from the web. Each face has been labeled with the name of the person pictured. 1680 of the people pictured have two or more distinct photos in the data set. The only constraint on these faces is that they were detected by the Viola-Jones face detector. More details can be found in the technical report below. Information: 13233 images 5749 people 1680 people with two or more images</description>
<link>https://academictorrents.com/download/692d556e6f2fcb430adeffc464eb8b0a6da58f65</link>
</item>
<item>
<title>Labeled Faces in the Wild (Dataset)</title>
<description>Welcome to Labeled Faces in the Wild, a database of face photographs designed for studying the problem of unconstrained face recognition. The data set contains more than 13,000 images of faces collected from the web. Each face has been labeled with the name of the person pictured. 1680 of the people pictured have two or more distinct photos in the data set. The only constraint on these faces is that they were detected by the Viola-Jones face detector. More details can be found in the technical report below. Information: 13233 images 5749 people 1680 people with two or more images Citation: Labeled Faces in the Wild: A Database for Studying Face Recognition in Unconstrained Environments Gary B. Huang and Manu Ramesh and Tamara Berg and Erik Learned-Miller University of Massachusetts, Amherst - 2007</description>
<link>https://academictorrents.com/download/9547ef95bc7007685afe52a8ec940aa61530bc99</link>
</item>
<item>
<title>The Street View House Numbers (SVHN) Dataset (Dataset)</title>
<description>SVHN is a real-world image dataset for developing machine learning and object recognition algorithms with minimal requirement on data preprocessing and formatting. It can be seen as similar in flavor to MNIST (e.g., the images are of small cropped digits), but incorporates an order of magnitude more labeled data (over 600,000 digit images) and comes from a significantly harder, unsolved, real world problem (recognizing digits and numbers in natural scene images). SVHN is obtained from house numbers in Google Street View images. Overview 10 classes, 1 for each digit. Digit  1  has label 1,  9  has label 9 and  0  has label 10. 73257 digits for training, 26032 digits for testing, and 531131 additional, somewhat less difficult samples, to use as extra training data Comes in two formats: 1. Original images with character level bounding boxes. 2. MNIST-like 32-by-32 images centered around a single character (many of the images do contain some distractors at the sides). These are the original, variable-resolution, color house-number images with character level bounding boxes, as shown in the examples images above. (The blue bounding boxes here are just for illustration purposes. The bounding box information are stored in digitStruct.mat instead of drawn directly on the images in the dataset.) Each tar.gz file contains the orignal images in png format, together with a digitStruct.mat file, which can be loaded using Matlab. The digitStruct.mat file contains a struct called digitStruct with the same length as the number of original images. Each element in digitStruct has the following fields: name which is a string containing the filename of the corresponding image. bbox which is a struct array that contains the position, size and label of each digit bounding box in the image. Eg: digitStruct(300).bbox(2).height gives height of the 2nd digit bounding box in the 300th image. Reference Please cite the following reference in papers using this dataset: Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, Andrew Y. Ng Reading Digits in Natural Images with Unsupervised Feature Learning NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011. Please use  as the URL for this site when necessary For questions regarding the dataset, please contact streetviewhousenumbers@gmail.com</description>
<link>https://academictorrents.com/download/6f4caf3c24803d114c3cae3ab9cb946cd23c7213</link>
</item>
<item>
<title>THE NORB DATASET, V1.0 (Dataset)</title>
<description>This database is intended for experiments in 3D object reocgnition from shape. It contains images of 50 toys belonging to 5 generic categories: four-legged animals, human figures, airplanes, trucks, and cars. The objects were imaged by two cameras under 6 lighting conditions, 9 elevations (30 to 70 degrees every 5 degrees), and 18 azimuths (0 to 340 every 20 degrees). The training set is composed of 5 instances of each category (instances 4, 6, 7, 8 and 9), and the test set of the remaining 5 instances (instances 0, 1, 2, 3, and 5). CONTENT The files are gzipped for download purpose. After uncompressed, they are in a simple binary matrix format, with file postfix ".mat". The file format is explained in a later section. The "-dat" files store the image sequences. The "-cat" files store the corresponding category of the images. Each "-dat" file stores 29,160 image pairs (6 categories, 5 instances, 6 lightings, 9 elevations, and 18 azimuths). The 6-th category is for images without objects, which can be used to train a system to reject images as none of the 5 object categories. Each corresponding "-cat" file contains 29,160 category labels (0 for animal, 1 for human, 2 for plane, 3 for truck, 4 for car, 5 for blank). Each "-info" file stores 29,160 10-dimensional vectors, which contain additional information about the corresponding images. The first 4 elements in the vector are: - 1. the instance in the category (0 to 9) - 2. the elevation (0 to 8, which mean cameras are 30, 35,40,45,50,55,60,65,70 degrees from the horizontal respectively) - 3. the azimuth (0,2,4,...,34, multiply by 10 to get the azimuth in degrees) - 4. the lighting condition (0 to 5) and the next 6 elements describe the peturbations added to the object when superposed onto a cluttered background. (see next section) For regular training and testing, "-dat" and "-cat" files are sufficient. "-info" files are provided in case some other forms of classification or preprocessing are needed. JITTERED OBJECTS AND CLUTTERED BACKGROUND After capturing, each image has been processed so that the object is centered in the image (the center of mass of object pixels are in the center of the image), scaled so that the bounding box is roughly 80x80 pixels, and placed on a uniform background, including the cast shadow. And then 3 sources of variations are added to the data set: - the objects are peturbed - the objects are superposed onto complex background - distractor objects are added to the background The objects are randomly peturbed in 5 ways. They are scaled by factors between 0.78 to 1.0; in-plane rotated -5 to +5 degrees; and shifted -6 to +6 pixels horizontally and vertically. The image intensities (in the range of 0 to 255) are a random value between -20 to +20; image contrasts are scaled in the range of 0.8 to 1.3. The peturbations are stored in the last 6 elements in the "-info" files: - 5. horizontal shift (-6 to +6) - 6. vertical shift (-6 to +6) - 7. lumination change (-20 to +20) - 8. contrast (0.8 to 1.3) - 9. object scale (0.78 to 1.0) - 10. rotation (-5 to +5 degrees) The complex background images are extracted from a subset of natural scene images from Corel image library. The images contain scenes with large region contrasts such as lake against moutain, and irregular region boundaries. One distractor object is added to each image. The distractor is located toward the boundary of the image, but can clutter the main object in the center. There are images with only background and distractor objects. These images belong to their own category, as indicated in the category files. FILE FORMAT The files are stored in the so-called "binary matrix" file format, which is a simple format for vectors and multidimensional matrices of various element types. Binary matrix files begin with a file header which describes the type and size of the matrix, and then comes the binary image of the matrix. The header is best described by a C structure: struct header  int magic; // 4 bytes int ndim; // 4 bytes, little endian int dim[3]; ; Note that when the matrix has less than 3 dimensions, say, it s a 1D vector, then dim[1] and dim[2] are both 1. When the matrix has more than 3 dimensions, the header will be followed by further dimension size information. Otherwise, after the file header comes the matrix data, which is stored with the index in the last dimension changes the fastest. The magic number encodes the element type of the matrix: - 0x1E3D4C51 for a single precision matrix - 0x1E3D4C52 for a packed matrix - 0x1E3D4C53 for a double precision matrix - 0x1E3D4C54 for an integer matrix - 0x1E3D4C55 for a byte matrix - 0x1E3D4C56 for a short matrix Since the files are generated on an Intel machine, they use the little-endian scheme to encode the 4-byte integers. Pay attention when you read the files on machines that use big-endian. - The "-dat" files store a 4D tensor of dimensions 29160x2x108x108. - The "-cat" files store a 1D vector of dimension 29,160. - The "-info" files store a 2D matrix of dimensions 29160x10. Here s a piece of Matlab code to show how to read some example files. (to avoid the endian confusion, we read bytes of the header):</description>
<link>https://academictorrents.com/download/4bcab1393bc699a39cb5c409e1f957d6c2918732</link>
</item>
<item>
<title>PASCAL-S - The Secrets of Salient Object Segmentation Dataset (Dataset)</title>
<description>Free-fiewing fixations on a subset of 850 images from PASCAL VOC.  Collected on 8 subjects, 3s viewing time, Eyelink II eye tracker. The performance of most algorithms suggest that PASCAL-S is less biased than most of the saliency datasets. 850 IMAGES FROM PASCAL 2010 1296 OBJECT INSTANCES 12 SUBJECTS     Folders in archive: algmaps/ algmaps/pascal algmaps/pascal/mcg_gbvs algmaps/pascal/humanFix algmaps/pascal/gc algmaps/pascal/dva algmaps/pascal/ft algmaps/pascal/sig algmaps/pascal/aim algmaps/pascal/pcas algmaps/pascal/gbvs algmaps/pascal/sun algmaps/pascal/aws algmaps/pascal/sf algmaps/pascal/itti algmaps/bruce algmaps/bruce/dva algmaps/bruce/sig algmaps/bruce/aim algmaps/bruce/gbvs algmaps/bruce/sun algmaps/bruce/aws algmaps/bruce/itti algmaps/cerf algmaps/cerf/dva algmaps/cerf/sig algmaps/cerf/aim algmaps/cerf/gbvs algmaps/cerf/sun algmaps/cerf/aws algmaps/cerf/itti algmaps/imgsal algmaps/imgsal/humanFix algmaps/imgsal/gc algmaps/imgsal/cpmc_gbvs algmaps/imgsal/dva algmaps/imgsal/ft algmaps/imgsal/sig algmaps/imgsal/aim algmaps/imgsal/pcas algmaps/imgsal/gbvs algmaps/imgsal/sun algmaps/imgsal/aws algmaps/imgsal/sf algmaps/imgsal/itti algmaps/ft algmaps/ft/gc algmaps/ft/cpmc_gbvs algmaps/ft/dva algmaps/ft/ft algmaps/ft/sig algmaps/ft/aim algmaps/ft/pcas algmaps/ft/gbvs algmaps/ft/sun algmaps/ft/aws algmaps/ft/sf algmaps/ft/itti algmaps/judd algmaps/judd/dva algmaps/judd/sig algmaps/judd/aim algmaps/judd/gbvs algmaps/judd/sun algmaps/judd/aws algmaps/judd/itti</description>
<link>https://academictorrents.com/download/6c49defd6f0e417c039637475cde638d1363037e</link>
</item>
<item>
<title>PASCAL-Context Dataset (Dataset)</title>
<description>This dataset is a set of additional annotations for PASCAL VOC 2010. It goes beyond the original PASCAL semantic segmentation task by providing annotations for the whole scene. The statistics section has a full list of 400+ labels. Every pixel has a unique class label. Instance information (i.e, different masks to separate different instances of the same class in the same image) are currently provided for the 20 PASCAL objects. Statistics Since the dataset is an annotation of PASCAL VOC 2010, it has the same statistics as those of the original dataset. Training and validation contains 10,103 images while testing contains 9,637 images. Usage Considerations The classes are not drawn from a fixed pool. Instead labelers were free to either select or type in what they believe to be the appropriate class and to determine what the appropriate object granularity is. We decided to merge/split some of the categories so the current number of categories is different from what we mentioned in the CVPR 2014 paper. When using this dataset it is important that you examine classes to ensure they match your intended use. For example, sand is often labeled independently despite also being considered ground. Those interested in ground may want to cluster sand and ground together along with other classes. Citation The Role of Context for Object Detection and Semantic Segmentation in the Wild Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, Alan Yuille IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014 Acknowledgements We would like to acknowledge the support by Implementation of Technologies for Identification, Behavior, and Location of Human based on Sensor Network Fusion Program through the Korean Ministry of Trade, Industry and Energy (Grant Number: 10041629). We would also like to thank National Science Foundation for grant 1317376 (Visual Cortex on Silicon. NSF Expedition in Computing). We thank Viet Nguyen for coordinating and leading the efforts for cleaning up the annotations.</description>
<link>https://academictorrents.com/download/eec6177ad62f4c47086e4cbec93ac4c08857ddbe</link>
</item>
<item>
<title>PASCAL-Part Dataset (Dataset)</title>
<description>This dataset is a set of additional annotations for PASCAL VOC 2010. It goes beyond the original PASCAL object detection task by providing segmentation masks for each body part of the object. For categories that do not have a consistent set of parts (e.g., boat), we provide the silhouette annotation. Statistics Since the dataset is an annotation of the PASCAL VOC 2010, it has the same statistics as those of the original dataset. Training and validation contains 10,103 images while testing contains 9,637 images. Usage Considerations We provide segmentation masks for detailed body parts. One can merge several parts to get appropriate object part granularity for different tasks. For instance, "eyes", "ears", "nose", etc. can be merged into a single "head" part. Citation Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts Xianjie Chen, Roozbeh Mottaghi, Xiaobai Liu, Sanja Fidler, Raquel Urtasun, Alan Yuille IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014 Acknowledgements We thank Viet Nguyen for coordinating and leading the efforts for cleaning up the annotations. We would like to acknowledge the support by grants ARO 62250-CS and N00014-12-1-0883.</description>
<link>https://academictorrents.com/download/f86670296bff85bcdffea6c4fc2e791446f9fb5e</link>
</item>
<item>
<title>Columbia University Image Library (COIL-20) (Dataset)</title>
<description>To database is available in two versions. The first, [unprocessed], consists of images for five of the objects that contain both the object and the background. The second, [processed], contains images for all of the objects in which the background has been discarded (and the images consist of the smallest square that contains the object). For formal documentation look at the corresponding compressed technical report "Columbia Object Image Library (COIL-20)," S. A. Nene, S. K. Nayar and H. Murase, Technical Report CUCS-005-96, February 1996.</description>
<link>https://academictorrents.com/download/1d16994c70b7fff8bfe917f83c397b1193daee7f</link>
</item>
<item>
<title>CIFAR-100 (Canadian Institute for Advanced Research) (Dataset)</title>
<description>This dataset is just like the CIFAR-10, except it has 100 classes containing 600 images each. There are 500 training images and 100 testing images per class. The 100 classes in the CIFAR-100 are grouped into 20 superclasses. Each image comes with a "fine" label (the class to which it belongs) and a "coarse" label (the superclass to which it belongs). Here is the list of classes in the CIFAR-100: SuperclassClasses aquatic mammalsbeaver, dolphin, otter, seal, whale fishaquarium fish, flatfish, ray, shark, trout flowersorchids, poppies, roses, sunflowers, tulips food containersbottles, bowls, cans, cups, plates fruit and vegetablesapples, mushrooms, oranges, pears, sweet peppers household electrical devicesclock, computer keyboard, lamp, telephone, television household furniturebed, chair, couch, table, wardrobe insectsbee, beetle, butterfly, caterpillar, cockroach large carnivoresbear, leopard, lion, tiger, wolf large man-made outdoor thingsbridge, castle, house, road, skyscraper large natural outdoor scenescloud, forest, mountain, plain, sea large omnivores and herbivorescamel, cattle, chimpanzee, elephant, kangaroo medium-sized mammalsfox, porcupine, possum, raccoon, skunk non-insect invertebratescrab, lobster, snail, spider, worm peoplebaby, boy, girl, man, woman reptilescrocodile, dinosaur, lizard, snake, turtle small mammalshamster, mouse, rabbit, shrew, squirrel treesmaple, oak, palm, pine, willow vehicles 1bicycle, bus, motorcycle, pickup truck, train vehicles 2lawn-mower, rocket, streetcar, tank, tractor Yes, I know mushrooms aren t really fruit or vegetables and bears aren t really carnivores.</description>
<link>https://academictorrents.com/download/9adb30144cf53809ec0613fa869b0a65b4e81ff5</link>
</item>
<item>
<title>Mnih Massachusetts Roads Dataset (Dataset)</title>
<description>"The datasets introduced in Chapter 6 of my PhD thesis are below. See the thesis for more details."</description>
<link>https://academictorrents.com/download/3b17f08ed5027ea24db04f460b7894d913f86c21</link>
</item>
</channel>
</rss>
