Sequence-to-Sequence learning using PyTorch

Last update: Nov 17, 2022

Overview

Seq2Seq in PyTorch

This is a complete suite for training sequence-to-sequence models in PyTorch. It consists of several models and code to both train and infer using them.

Using this code you can train:

Neural-machine-translation (NMT) models
Language models
Image to caption generation
Skip-thought sentence representations
And more...

Installation

git clone --recursive https://github.com/eladhoffer/seq2seq.pytorch
cd seq2seq.pytorch; python setup.py develop

Models

Models currently available:

Simple Seq2Seq recurrent model
Recurrent Seq2Seq with attentional decoder
Google neural machine translation (GNMT) recurrent model
Transformer - attention-only model from "Attention Is All You Need"

Datasets

Datasets currently available:

WMT16
WMT17
OpenSubtitles 2016
COCO image captions
Conceptual captions

All datasets can be tokenized using 3 available segmentation methods:

Character based segmentation
Word based segmentation
Byte-pair-encoding (BPE) as suggested by bpe with selectable number of tokens.

After choosing a tokenization method, a vocabulary will be generated and saved for future inference.

Training methods

The models can be trained using several methods:

Basic Seq2Seq - given encoded sequence, generate (decode) output sequence. Training is done with teacher-forcing.
Multi Seq2Seq - where several tasks (such as multiple languages) are trained simultaneously by using the data sequences as both input to the encoder and output for decoder.
Image2Seq - used to train image to caption generators.

Usage

Example training scripts are available in scripts folder. Inference examples are available in examples folder.

example for training a transformer on WMT16 according to original paper regime:

DATASET=${1:-"WMT16_de_en"}
DATASET_DIR=${2:-"./data/wmt16_de_en"}
OUTPUT_DIR=${3:-"./results"}

WARMUP="4000"
LR0="512**(-0.5)"

python main.py \
  --save transformer \
  --dataset ${DATASET} \
  --dataset-dir ${DATASET_DIR} \
  --results-dir ${OUTPUT_DIR} \
  --model Transformer \
  --model-config "{'num_layers': 6, 'hidden_size': 512, 'num_heads': 8, 'inner_linear': 2048}" \
  --data-config "{'moses_pretok': True, 'tokenization':'bpe', 'num_symbols':32000, 'shared_vocab':True}" \
  --b 128 \
  --max-length 100 \
  --device-ids 0 \
  --label-smoothing 0.1 \
  --trainer Seq2SeqTrainer \
  --optimization-config "[{'step_lambda':
                          \"lambda t: { \
                              'optimizer': 'Adam', \
                              'lr': ${LR0} * min(t ** -0.5, t * ${WARMUP} ** -1.5), \
                              'betas': (0.9, 0.98), 'eps':1e-9}\"
                          }]"

example for training attentional LSTM based model with 3 layers in both encoder and decoder:

python main.py \
  --save de_en_wmt17 \
  --dataset ${DATASET} \
  --dataset-dir ${DATASET_DIR} \
  --results-dir ${OUTPUT_DIR} \
  --model RecurrentAttentionSeq2Seq \
  --model-config "{'hidden_size': 512, 'dropout': 0.2, \
                   'tie_embedding': True, 'transfer_hidden': False, \
                   'encoder': {'num_layers': 3, 'bidirectional': True, 'num_bidirectional': 1, 'context_transform': 512}, \
                   'decoder': {'num_layers': 3, 'concat_attention': True,\
                               'attention': {'mode': 'dot_prod', 'dropout': 0, 'output_transform': True, 'output_nonlinearity': 'relu'}}}" \
  --data-config "{'moses_pretok': True, 'tokenization':'bpe', 'num_symbols':32000, 'shared_vocab':True}" \
  --b 128 \
  --max-length 80 \
  --device-ids 0 \
  --trainer Seq2SeqTrainer \
  --optimization-config "[{'epoch': 0, 'optimizer': 'Adam', 'lr': 1e-3},
                          {'epoch': 6, 'lr': 5e-4},
                          {'epoch': 8, 'lr':1e-4},
                          {'epoch': 10, 'lr': 5e-5},
                          {'epoch': 12, 'lr': 1e-5}]" \

Sequence-to-Sequence learning using PyTorch

Related tags

Overview

Seq2Seq in PyTorch

Installation

Models

Datasets

Training methods

Usage

Owner

Elad Hoffer

Universal Adversarial Triggers for Attacking and Analyzing NLP (EMNLP 2019)

Easy, fast, effective, and automatic g-code compression!

AI-powered literature discovery and review engine for medical/scientific papers

Code for CVPR 2021 paper: Revamping Cross-Modal Recipe Retrieval with Hierarchical Transformers and Self-supervised Learning

Control the classic General Instrument SP0256-AL2 speech chip and AY-3-8910 sound generator with a Raspberry Pi and this Python library.

A curated list of efficient attention modules

An easy to use, user-friendly and efficient code for extracting OpenAI CLIP (Global/Grid) features from image and text respectively.

Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing

Contact Extraction with Question Answering.

Beyond Paragraphs: NLP for Long Sequences

🧪 Cutting-edge experimental spaCy components and features

GPT-2 Model for Leetcode Questions in python

DeepAmandine is an artificial intelligence that allows you to talk to it for hours, you won't know the difference.

自然言語で書かれた時間情報表現を抽出/規格化するルールベースの解析器

Sinkhorn Transformer - Practical implementation of Sparse Sinkhorn Attention

Lightweight utility tools for the detection of multiple spellings, meanings, and language-specific terminology in British and American English

내부 작업용 django + vue(vuetify) boilerplate. 짠 하면 돌아감.

BMInf (Big Model Inference) is a low-resource inference package for large-scale pretrained language models (PLMs).

Bidirectional Variational Inference for Non-Autoregressive Text-to-Speech (BVAE-TTS)

Deduplication is the task to combine different representations of the same real world entity.