This is the official PyTorch implementation of the paper "TransFG: A Transformer Architecture for Fine-grained Recognition" (Ju He, Jie-Neng Chen, Shuai Liu, Adam Kortylewski, Cheng Yang, Yutong Bai, Changhu Wang, Alan Yuille).

Last update: Jan 03, 2023

Overview

TransFG: A Transformer Architecture for Fine-grained Recognition

Official PyTorch code for the paper: TransFG: A Transformer Architecture for Fine-grained Recognition

Implementation based on DeiT pretrained on ImageNet-1K with distillation fine-tuning will be released soon.

Framework

Dependencies:

Python 3.7.3
PyTorch 1.5.1
torchvision 0.6.1
ml_collections

Usage

1. Download Google pre-trained ViT models

Get models in this link: ViT-B_16, ViT-B_32...

wget https://storage.googleapis.com/vit_models/imagenet21k/{MODEL_NAME}.npz

2. Prepare data

In the paper, we use data from 5 publicly available datasets:

Please download them from the official websites and put them in the corresponding folders.

3. Install required packages

Install dependencies with the following command:

pip3 install -r requirements.txt

4. Train

To train TransFG on CUB-200-2011 dataset with 4 gpus in FP-16 mode for 10000 steps run:

CUDA_VISIBLE_DEVICES=0,1,2,3 python3 -m torch.distributed.launch --nproc_per_node=4 train.py --dataset CUB_200_2011 --split overlap --num_steps 10000 --fp16 --name sample_run

Citation

If you find our work helpful in your research, please cite it as:

@article{he2021transfg,
  title={TransFG: A Transformer Architecture for Fine-grained Recognition},
  author={He, Ju and Chen, Jieneng and Liu, Shuai and Kortylewski, Adam and Yang, Cheng and Bai, Yutong and Wang, Changhu and Yuille, Alan},
  journal={arXiv preprint arXiv:2103.07976},
  year={2021}
}

Acknowledgement

Many thanks to ViT-pytorch for the PyTorch reimplementation of An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

This is the official PyTorch implementation of the paper "TransFG: A Transformer Architecture for Fine-grained Recognition" (Ju He, Jie-Neng Chen, Shuai Liu, Adam Kortylewski, Cheng Yang, Yutong Bai, Changhu Wang, Alan Yuille).

Related tags

Overview

TransFG: A Transformer Architecture for Fine-grained Recognition

Framework

Dependencies:

Usage

1. Download Google pre-trained ViT models

2. Prepare data

3. Install required packages

4. Train

Citation

Acknowledgement

Owner

Ju He

OCR software for recognition of handwritten text

Repository for Scene Text Detection with Supervised Pyramid Context Network with tensorflow.

text detection mainly based on ctpn model in tensorflow, id card detect, connectionist text proposal network

Erosion and dialation using structure element in OpenCV python

Line based ATR Engine based on OCRopy

ISI's Optical Character Recognition (OCR) software for machine-print and handwriting data

Official implementation of "An Image is Worth 16x16 Words, What is a Video Worth?" (2021 paper)

Detect textlines in document images

deployment of a hybrid model for automatic weapon detection/ anomaly detection for surveillance applications

Geometric Augmentation for Text Image

Fine tuning keras-ocr python package with custom synthetic dataset from scratch

TextField: Learning A Deep Direction Field for Irregular Scene Text Detection (TIP 2019)

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

fishington.io bot with OpenCV and NumPy

Regions sanitàries (RS), Sectors Sanitàris (SS) i Àrees Bàsiques de Salut (ABS) de Catalunya

Text language identification using Wikipedia data

CNN+LSTM+CTC based OCR implemented using tensorflow.

This is a GUI for scrapping PDFs with the help of optical character recognition making easier than ever to scrape PDFs.

Text-to-Image generation

Automatically fishes for you while you are afk :)