Code for the head detector (HeadHunter) proposed in our CVPR 2021 paper Tracking Pedestrian Heads in Dense Crowd.

Last update: Dec 06, 2022

Related tags

Overview

Head Detector

Code for the head detector (HeadHunter) proposed in our CVPR 2021 paper Tracking Pedestrian Heads in Dense Crowd. The head_detection module can be installed using pip in order to be able to plug-and-play with HeadHunter-T.

Requirements

Nvidia Driver >= 418
Cuda 10.0 and compaitible CudNN
Python packages : To install the required python packages; conda env create -f head_detection.yml.
Use the anaconda environment head_detection by activating it, source activate head_detection or conda activate head_detection.
Alternatively pip can be used to install required packages using pip install -r requirements.txt or update your existing environment with the aforementioned yml file.

Training

To train a model, define environment variable NGPU, config file and use the following command

$python -m torch.distributed.launch --nproc_per_node=$NGPU --use_env train.py --cfg_file config/config_chuman.yaml --world_size $NGPU --num_workers 4

Training is currently supported over (a) ScutHead dataset (b) CrowdHuman + ScutHead combined, (c) Our proposed CroHD dataset. This can be mentioned in the config file.
To train the model, config files must be defined. More details about the config files are mentioned in the section below

Evaluation and Testing

Unlike the training, testing and evaluation does not have a config file. Rather, all the parameters are set as argument variable while executing the code. Refer to the respective files, evaluate.py and test.py.
evaluate.py evaluates over the validation/test set using AP, MMR, F1, MODA and MODP metrics.
test.py runs the detector over a "bunch of images" in the testing set for qualitative evaluation.

Config file

A config file is necessary for all training. It's built to ease the number of arg variable passed during each execution. Each sub-sections are as elaborated below.

DATASET
1. Set the base_path as the parent directory where the dataset is situated at.
2. Train and Valid are .txt files that contains relative path to respective images from the base_path defined above and their corresponding Ground Truth in (x_min, y_min, x_max, y_max) format. Generation files for the three datasets can be seen inside data directory. For example,
```
/path/to/image.png
x_min_1, y_min_1, x_max_1, y_max_1
x_min_2, y_min_2, x_max_2, y_max_2
x_min_3, y_min_3, x_max_3, y_max_3
.
.
.
```
1. mean_std are RGB means and stdev of the training dataset. If not provided, can be computed prior to the start of the training
TRAINING
1. Provide pretrained_model and corresponding start_epoch for resuming.
2. milestones are epoch at which the learning rates are set to 0.1 * lr.
3. only_backbone option loads just the Resnet backbone and not the head. Not applicable for mobilenet.
NETWORK
1. The mentioned parameters are as described in experiment section of the paper.
2. When using median_anchors, the anchors have to be defined in anchors.py.
3. We experimented with mobilenet, resnet50 and resnet150 as alternative backbones. This experiment was not reported in the paper due to space constraints. We found the accuracy to significantly decrease with mobilenet but resnet50 and resnet150 yielded an almost same performance.
4. We also briefly experimented with Deformable Convolutions but again didn't see noticable improvements in performance. The code we used are available in this repository.

Note :

This codebase borrows a noteable portion from pytorch-vision owing to the fact some of their modules cannot be "imported" as a package.

Citation :

@InProceedings{Sundararaman_2021_CVPR,
    author    = {Sundararaman, Ramana and De Almeida Braga, Cedric and Marchand, Eric and Pettre, Julien},
    title     = {Tracking Pedestrian Heads in Dense Crowd},
    booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    month     = {June},
    year      = {2021},
    pages     = {3865-3875}
}

Code for the head detector (HeadHunter) proposed in our CVPR 2021 paper Tracking Pedestrian Heads in Dense Crowd.

Related tags

Overview

Head Detector

Requirements

Training

Evaluation and Testing

Config file

Note :

Citation :

Owner

Ramana Subramanyam

A program that takes in the hand gesture displayed by the user and translates ASL.

computer vision, image processing and machine learning on the web browser or node.

Deskew is a command line tool for deskewing scanned text documents. It uses Hough transform to detect "text lines" in the image. As an output, you get an image rotated so that the lines are horizontal.

Select range and every time the screen changes, OCR is activated.

Introduction to Augmented Reality (AR) with Python 3 and OpenCV 4.2.

A semi-automatic open-source tool for Layout Analysis and Region EXtraction on early printed books.

Textboxes_plusplus implementation with Tensorflow (python)

A small C++ implementation of LSTM networks, focused on OCR.

Rotational region detection based on Faster-RCNN.

python ocr using tesseract/ with EAST opencv detector

This is a GUI for scrapping PDFs with the help of optical character recognition making easier than ever to scrape PDFs.

Repository collecting all the submodules for the new PyTorch-based OCR System.

Um simples projeto para fazer o reconhecimento do captcha usado pelo jogo bombcrypto

Script para controlar o movimento do mouse usando Python e openCV com câmera em tempo real que detecta pontos de referência da mão, rastreia padrões de gestos em vez de um mouse físico.

Implement 'Single Shot Text Detector with Regional Attention, ICCV 2017 Spotlight'

Slice a single image into multiple pieces and create a dataset from them

Let's explore how we can extract text from forms

Hiiii this is the Spanish for Linux and win 10 and in the near future the english version of PortScan my new tool on which you can see what ports are Open only with the IP adress.

Ocular is a state-of-the-art historical OCR system.

Implementation of our paper 'PixelLink: Detecting Scene Text via Instance Segmentation' in AAAI2018