ML-Decoder: Scalable and Versatile Classification Head

Last update: Jan 04, 2023

Related tags

Deep Learning ML_Decoder

Overview

ML-Decoder: Scalable and Versatile Classification Head

Paper

Official PyTorch Implementation

Tal Ridnik, Gilad Sharir, Avi Ben-Cohen, Emanuel Ben-Baruch, Asaf Noy
DAMO Academy, Alibaba Group

Abstract

In this paper, we introduce ML-Decoder, a new attention-based classification head. ML-Decoder predicts the existence of class labels via queries, and enables better utilization of spatial data compared to global average pooling. By redesigning the decoder architecture, and using a novel group-decoding scheme, ML-Decoder is highly efficient, and can scale well to thousands of classes. Compared to using a larger backbone, ML-Decoder consistently provides a better speed-accuracy trade-off. ML-Decoder is also versatile - it can be used as a drop-in replacement for various classification heads, and generalize to unseen classes when operated with word queries. Novel query augmentations further improve its generalization ability. Using ML-Decoder, we achieve state-of-the-art results on several classification tasks: on MS-COCO multi-label, we reach 91.4% mAP; on NUS-WIDE zero-shot, we reach 31.1% ZSL mAP; and on ImageNet single-label, we reach with vanilla ResNet50 backbone a new top score of 80.7%, without extra data or distillation.

ML-Decoder Implementation

ML-Decoder implementation is available here. It can be easily integrated into any backbone using this example code:

ml_decoder_head = MLDecoder(num_classes) # initilization

spatial_embeddings = self.backbone(input_image) # backbone generates spatial embeddings      
 
logits = ml_decoder_head(spatial_embeddings) # transfrom spatial embeddings to logits

Training Code

We will share a full reproduction code for the article results.

Multi-label Training Code

A reproduction code for MS-COCO multi-label:

python train.py  \
--data=/home/datasets/coco2014/ \
--model_name=tresnet_l \
--image_size=448

Single-label Training Code

Our single-label training code uses the excellent timm repo. Reproduction code is currently from a fork, we will work toward a full merge to the main repo.

git clone https://github.com/mrT23/pytorch-image-models.git

This is the code for A2 configuration training, with ML-Decoder (--use-ml-decoder-head=1):

python -u -m torch.distributed.launch --nproc_per_node=8 \
--nnodes=1 \
--node_rank=0 \
./train.py \
/data/imagenet/ \
--amp \
-b=256 \
--epochs=300 \
--drop-path=0.05 \
--opt=lamb \
--weight-decay=0.02 \
--sched='cosine' \
--lr=4e-3 \
--warmup-epochs=5 \
--model=resnet50 \
--aa=rand-m7-mstd0.5-inc1 \
--reprob=0.0 \
--remode='pixel' \
--mixup=0.1 \
--cutmix=1.0 \
--aug-repeats 3 \
--bce-target-thresh 0.2 \
--smoothing=0 \
--bce-loss \
--train-interpolation=bicubic \
--use-ml-decoder-head=1

ZSL Training Code

Reproduction code for ZSL is WIP.

Citation

@misc{ridnik2021mldecoder,
      title={ML-Decoder: Scalable and Versatile Classification Head}, 
      author={Tal Ridnik and Gilad Sharir and Avi Ben-Cohen and Emanuel Ben-Baruch and Asaf Noy},
      year={2021},
      eprint={2111.12933},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}

ML-Decoder: Scalable and Versatile Classification Head

Related tags

Overview

ML-Decoder: Scalable and Versatile Classification Head

ML-Decoder Implementation

Training Code

Multi-label Training Code

Single-label Training Code

ZSL Training Code

Citation

Owner

This repo contains the official code of our work SAM-SLR which won the CVPR 2021 Challenge on Large Scale Signer Independent Isolated Sign Language Recognition.

An executor that performs image segmentation on fashion items

Modeling Category-Selective Cortical Regions with Topographic Variational Autoencoders

Semi-SDP Semi-supervised parser for semantic dependency parsing.

Türkiye Canlı Mobese Görüntülerinde Profesyonel Nesne Takip Sistemi

Classification of EEG data using Deep Learning

Safe Policy Optimization with Local Features

Deep Residual Learning for Image Recognition

OMNIVORE is a single vision model for many different visual modalities

Real-ESRGAN aims at developing Practical Algorithms for General Image Restoration.

A curated list of automated deep learning (including neural architecture search and hyper-parameter optimization) resources.

[IROS2021] NYU-VPR: Long-Term Visual Place Recognition Benchmark with View Direction and Data Anonymization Influences

Blender scripts for computing geodesic distance

Unofficial implementation of Proxy Anchor Loss for Deep Metric Learning

Code repository for the paper Computer Vision User Entity Behavior Analytics

Simple Linear 2nd ODE Solver GUI - A 2nd constant coefficient linear ODE solver with simple GUI using euler's method

Protect against subdomain takeover

Robust, modular and efficient implementation of advanced Hamiltonian Monte Carlo algorithms

Vrcwatch - Supply the local time to VRChat as Avatar Parameters through OSC

Fight Recognition from Still Images in the Wild @ WACVW2022, Real-world Surveillance Workshop