IndoBERTweet is the first large-scale pretrained model for Indonesian Twitter. Published at EMNLP 2021 (main conference)

Last update: Nov 30, 2022

Overview

IndoBERTweet 🐦 🇮🇩

1. Paper

Fajri Koto, Jey Han Lau, and Timothy Baldwin. IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021), Dominican Republic (virtual).

2. About

IndoBERTweet is the first large-scale pretrained model for Indonesian Twitter that is trained by extending a monolingually trained Indonesian BERT model with additive domain-specific vocabulary.

In this paper, we show that initializing domain-specific vocabulary with average-pooling of BERT subword embeddings is more efficient than pretraining from scratch, and more effective than initializing based on word2vec projections.

3. Pretraining Data

We crawl Indonesian tweets over a 1-year period using the official Twitter API, from December 2019 to December 2020, with 60 keywords covering 4 main topics: economy, health, education, and government. We obtain in total of 409M word tokens, two times larger than the training data used to pretrain IndoBERT. Due to Twitter policy, this pretraining data will not be released to public.

4. How to use

Load model and tokenizer (tested with transformers==3.5.1)

from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("indolem/indobertweet-base-uncased")
model = AutoModel.from_pretrained("indolem/indobertweet-base-uncased")

Preprocessing Steps:

lower-case all words
converting user mentions and URLs into @USER and HTTPURL, respectively
translating emoticons into text using the emoji package.

5. Results over 7 Indonesian Twitter Datasets

Models	Sentiment		Emotion	Hate Speech		NER		Average
Models	IndoLEM	SmSA	EmoT	HS1	HS2	Formal	Informal	Average
mBERT	76.6	84.7	67.5	85.1	75.1	85.2	83.2	79.6
malayBERT	82.0	84.1	74.2	85.0	81.9	81.9	81.3	81.5
IndoBERT (Willie, et al., 2020)	84.1	88.7	73.3	86.8	80.4	86.3	84.3	83.4
IndoBERT (Koto, et al., 2020)	84.1	87.9	71.0	86.4	79.3	88.0	86.9	83.4
IndoBERTweet (1M steps from scratch)	86.2	90.4	76.0	88.8	87.5	88.1	85.4	86.1
IndoBERT + Voc adaptation + 200k steps	86.6	92.7	79.0	88.4	84.0	87.7	86.9	86.5

IndoBERTweet is the first large-scale pretrained model for Indonesian Twitter. Published at EMNLP 2021 (main conference)

Related tags

Overview

IndoBERTweet 🐦 🇮🇩

1. Paper

2. About

3. Pretraining Data

4. How to use

5. Results over 7 Indonesian Twitter Datasets

Owner

IndoLEM

Create a machine learning model which will predict if the mortgage will be approved or not based on 5 variables

This project aims to conduct a text information retrieval and text mining on medical research publication regarding Covid19 - treatments and vaccinations.

Contract Understanding Atticus Dataset

Study German declensions (dER nettE Mann, ein nettER Mann, mit dEM nettEN Mann, ohne dEN nettEN Mann ...) Generate as many exercises as you want using the incredible power of SPACY!

NLPIR tutorial: pretrain for IR. pre-train on raw textual corpus, fine-tune on MS MARCO Document Ranking

Analyse japanese ebooks using MeCab to determine the difficulty level for japanese learners

A Fast Sequence Transducer Implementation with PyTorch Bindings

Implementaion of our ACL 2022 paper Bridging the Data Gap between Training and Inference for Unsupervised Neural Machine Translation

Prompt-learning is the latest paradigm to adapt pre-trained language models (PLMs) to downstream NLP tasks

基于“Seq2Seq+前缀树”的知识图谱问答

To be a next-generation DL-based phenotype prediction from genome mutations.

This is the main repository of open-sourced speech technology by Huawei Noah's Ark Lab.

Paddlespeech Streaming ASR GUI

Repository for Graph2Pix: A Graph-Based Image to Image Translation Framework

This project uses word frequency and Term Frequency-Inverse Document Frequency to summarize a text.

TalkNet: Audio-visual active speaker detection Model

A collection of Classical Chinese natural language processing models, including Classical Chinese related models and resources on the Internet.

Unofficial PyTorch implementation of Google AI's VoiceFilter system

TEACh is a dataset of human-human interactive dialogues to complete tasks in a simulated household environment.

Uses Google's gTTS module to easily create robo text readin' on command.