This repository has a implementations of data augmentation for NLP for Japanese.

Last update: Nov 11, 2022

Related tags

Text Data & NLP daaja

Overview

daaja

This repository has a implementations of data augmentation for NLP for Japanese:

EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks
An Analysis of Simple Data Augmentation for Named Entity Recognition

Install

pip install daaja

How to use

EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks

Command

python -m aug_ja.eda.run --input input.tsv --output data_augmentor.tsv

The format of input.tsv is as follows:

1	この映画はとてもおもしろい
0	つまらない映画だった

In Python

from aug_ja.eda import EasyDataAugmentor
augmentor = EasyDataAugmentor(alpha_sr=0.1, alpha_ri=0.1, alpha_rs=0.1, p_rd=0.1, num_aug=4)
text = "日本語でデータ拡張を行う"
aug_texts = augmentor.augments(text)
print(aug_texts)
# ['日本語でを拡張データ行う', '日本語でデータ押広げるを行う', '日本語でデータ拡張を行う', '日本語で智見拡張を行う', '日本語でデータ拡張を行う']

An Analysis of Simple Data Augmentation for Named Entity Recognition

Command

python -m aug_ja.ner_sda.run --input input.tsv --output data_augmentor.tsv

The format of input.tsv is as follows:

私	O
は	O
田中	B-PER
と	O
いい	O
ます	O

In Python

from daaja.ner_sda import SimpleDataAugmentationforNER
tokens_list = [
    ["私", "は", "田中", "と", "いい", "ます"],
    ["筑波", "大学", "に", "所属", "して", "ます"],
    ["今日", "から", "筑波", "大学", "に", "通う"],
    ["茨城", "大学"],
]
labels_list = [
    ["O", "O", "B-PER", "O", "O", "O"],
    ["B-ORG", "I-ORG", "O", "O", "O", "O"],
    ["B-DATE", "O", "B-ORG", "I-ORG", "O", "O"],
    ["B-ORG", "I-ORG"],
]
augmentor = SimpleDataAugmentationforNER(tokens_list=tokens_list, labels_list=labels_list,
                                            p_power=1, p_lwtr=1, p_mr=1, p_sis=1, p_sr=1, num_aug=4)
tokens = ["吉田", "さん", "は", "株式", "会社", "A", "に", "出張", "予定", "だ"]
labels = ["B-PER", "O", "O", "B-ORG", "I-ORG", "I-ORG", "O", "O", "O", "O"]
augmented_tokens_list, augmented_labels_list = augmentor.augments(tokens, labels)
print(augmented_tokens_list)
# [['吉田', 'さん', 'は', '株式', '会社', 'A', 'に', '出張', '志す', 'だ'],
#  ['吉田', 'さん', 'は', '株式', '大学', '大学', 'に', '出張', '予定', 'だ'],
#  ['吉田', 'さん', 'は', '株式', '会社', 'A', 'に', '出張', '予定', 'だ'],
#  ['吉田', 'さん', 'は', '筑波', '大学', 'に', '出張', '予定', 'だ'],
#  ['吉田', 'さん', 'は', '株式', '会社', 'A', 'に', '出張', '予定', 'だ']]
print(augmented_labels_list)
# [['B-PER', 'O', 'O', 'B-ORG', 'I-ORG', 'I-ORG', 'O', 'O', 'O', 'O'],
#  ['B-PER', 'O', 'O', 'B-ORG', 'I-ORG', 'I-ORG', 'O', 'O', 'O', 'O'],
#  ['B-PER', 'O', 'O', 'B-ORG', 'I-ORG', 'I-ORG', 'O', 'O', 'O', 'O'],
#  ['B-PER', 'O', 'O', 'B-ORG', 'I-ORG', 'O', 'O', 'O', 'O'],
#  ['B-PER', 'O', 'O', 'B-ORG', 'I-ORG', 'I-ORG', 'O', 'O', 'O', 'O']]

Reference

Comments

too many progress bars

When I use EasyDataAugmentor in the train process, there are too many progress bars in the console.

So, can you make this line 19 tqdm selectable on-off when we define EasyDataAugmentor? https://github.com/kajyuuen/daaja/blob/12835943868d43f5c248cf1ea87ab60f67a6e03d/daaja/flows/sequential_flow.py#L19

opened by Yongtae723 6
from daaja.methods.eda.easy_data_augmentor import EasyDataAugmentorにてエラー

daajaをpipインストール後、from daaja.methods.eda.easy_data_augmentor import EasyDataAugmentorを行うと、以下のエラーとなる。 ConnectionError: HTTPConnectionPool(host='compling.hss.ntu.edu.sg', port=80): Max retries exceeded with url: /wnja/data/1.1/wnjpn.db.gz (Caused by NewConnectionError('<urllib3.connection.HTTPConnection object at 0x7f3b6a6cced0>: Failed to establish a new connection: [Errno 110] Connection timed out'))

opened by naoki1213mj 5
is it possible to use on GPU device?

Hi!

thank you for the great library. when I train with this augmentation, this takes so much more time than forward and backward process.

therefore, can we possibly use this augmentation on GPU to save time?

thank you

opened by Yongtae723 3
Bump joblib from 1.1.0 to 1.2.0
Bumps joblib from 1.1.0 to 1.2.0.

Changelog

Sourced from joblib's changelog.

Release 1.2.0

Fix a security issue where eval(pre_dispatch) could potentially run arbitrary code. Now only basic numerics are supported. joblib/joblib#1327

Make sure that joblib works even when multiprocessing is not available, for instance with Pyodide joblib/joblib#1256

Avoid unnecessary warnings when workers and main process delete the temporary memmap folder contents concurrently. joblib/joblib#1263

Fix memory alignment bug for pickles containing numpy arrays. This is especially important when loading the pickle with mmap_mode != None as the resulting numpy.memmap object would not be able to correct the misalignment without performing a memory copy. This bug would cause invalid computation and segmentation faults with native code that would directly access the underlying data buffer of a numpy array, for instance C/C++/Cython code compiled with older GCC versions or some old OpenBLAS written in platform specific assembly. joblib/joblib#1254

Vendor cloudpickle 2.2.0 which adds support for PyPy 3.8+.

Vendor loky 3.3.0 which fixes several bugs including:

robustly forcibly terminating worker processes in case of a crash (joblib/joblib#1269);

avoiding leaking worker processes in case of nested loky parallel calls;

reliability spawn the correct number of reusable workers.

Release 1.1.1

Fix a security issue where eval(pre_dispatch) could potentially run arbitrary code. Now only basic numerics are supported. joblib/joblib#1327

Commits

5991350 Release 1.2.0

3fa2188 MAINT cleanup numpy warnings related to np.matrix in tests (#1340)

cea26ff CI test the future loky-3.3.0 branch (#1338)

8aca6f4 MAINT: remove pytest.warns(None) warnings in pytest 7 (#1264)

067ed4f XFAIL test_child_raises_parent_exits_cleanly with multiprocessing (#1339)

ac4ebd5 MAINT add back pytest warnings plugin (#1337)

a23427d Test child raises parent exits cleanly more reliable on macos (#1335)

ac09691 [MAINT] various test updates (#1334)

4a314b1 Vendor loky 3.2.0 (#1333)

bdf47e9 Make test_parallel_with_interactively_defined_functions_default_backend timeo...

Additional commits viewable in compare view

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.

Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

@dependabot rebase will rebase this PR

@dependabot recreate will recreate this PR, overwriting any edits that have been made to it

@dependabot merge will merge this PR after your CI passes on it

@dependabot squash and merge will squash and merge this PR after your CI passes on it

@dependabot cancel merge will cancel a previously requested merge and block automerging

@dependabot reopen will reopen this PR if it is closed

@dependabot close will close this PR and stop Dependabot recreating it. You can achieve the same result by closing it manually

@dependabot ignore this major version will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)

@dependabot ignore this minor version will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)

@dependabot ignore this dependency will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)

@dependabot use these labels will set the current labels as the default for future PRs for this repo and language

@dependabot use these reviewers will set the current reviewers as the default for future PRs for this repo and language

@dependabot use these assignees will set the current assignees as the default for future PRs for this repo and language

@dependabot use this milestone will set the current milestone as the default for future PRs for this repo and language

You can disable automated security fix PRs for this repo from the Security Alerts page.

dependencies
opened by dependabot[bot] 0
Implement Data Augmentation using Pre-trained Transformer Models
paper

Data Augmentation using Pre-trained Transformer Models

code

https://github.com/varunkumar-dev/TransformersDataAugmentation

ref

https://www.ai-shift.co.jp/techblog/1939

add-new-technique
opened by kajyuuen 0
Implement Contextual Augmentation
Paper

Contextual Augmentation: Data Augmentation by Words with Paradigmatic Relations

Code

https://github.com/pfnet-research/contextual_augmentation

add-new-technique
opened by kajyuuen 0
Implement MixText
Paper

MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text Classification

Code

https://github.com/GT-SALT/MixText

add-new-technique
opened by kajyuuen 0

Releases(v0.0.7)

v0.0.7(Oct 24, 2022)
Changes

Change pytest @kajyuuen (#35 #37 #38)

Change WORDNER_URL @kajyuuen (#34)

Source code(tar.gz)
Source code(zip)
daaja-0.0.7-py3-none-any.whl(18.19 KB)
v0.0.6(Mar 3, 2022)
Changes

Update version @kajyuuen (#27)

Add verbose option @kajyuuen (#25)

📖 Documentation

Add README_ja.md and Update README.md @kajyuuen (#26)

Source code(tar.gz)
Source code(zip)
v0.0.5(Feb 27, 2022)
Changes

💪 Enhancement

Add ContextualAugmentor @kajyuuen (#23)

Add BackTranslationAugmentor @kajyuuen (#21 , #22)

📖 Documentation

Add quick_example @kajyuuen (#17)

Source code(tar.gz)
Source code(zip)
v0.0.4(Feb 21, 2022)
Changes

Release v0.0.4 @kajyuuen (#16)

Chore add release drafter @kajyuuen (#6)

💪 Enhancement

Add tqdm @kajyuuen (#8)

📖 Documentation

Refactoring @kajyuuen (#15)

Add SDA example @kajyuuen (#9)

Add EDA example @kajyuuen (#7)

Source code(tar.gz)
Source code(zip)
v0.0.3(Feb 13, 2022)

Source code(tar.gz)
Source code(zip)
daaja-0.0.3-py3-none-any.whl(14.80 KB)
v0.0.2(Feb 13, 2022)

Source code(tar.gz)
Source code(zip)
daaja-0.0.2-py3-none-any.whl(14.97 KB)

Owner

Koga Kobayashi

GitHub Repository

Enterprise Scale NLP with Hugging Face & SageMaker Workshop series

Workshop: Enterprise-Scale NLP with Hugging Face & Amazon SageMaker Earlier this year we announced a strategic collaboration with Amazon to make it ea

161 Dec 16, 2022

File-based TF-IDF: Calculates keywords in a document, using a word corpus.

File-based TF-IDF Calculates keywords in a document, using a word corpus. Why? Because I found myself with hundreds of plain text files, with no way t

1 Feb 11, 2022

LeBenchmark: a reproducible framework for assessing SSL from speech

11 Nov 30, 2022

A list of NLP(Natural Language Processing) tutorials

NLP Tutorial A list of NLP(Natural Language Processing) tutorials built on PyTorch. Table of Contents A step-by-step tutorial on how to implement and

1.3k Dec 25, 2022

Open-source offline translation library written in Python. Uses OpenNMT for translations

Open source neural machine translation in Python. Designed to be used either as a Python library or desktop application. Uses OpenNMT for translations and PyQt for GUI.

1.6k Jan 01, 2023

ChessCoach is a neural network-based chess engine capable of natural-language commentary.

380 Dec 03, 2022

justCTF [*] 2020 challenges sources

justCTF [*] 2020 This repo contains sources for justCTF [*] 2020 challenges hosted by justCatTheFish. TLDR: Run a challenge with ./run.sh (requires Do

25 Dec 27, 2022

Input english text, then translate it between languages n times using the Deep Translator Python Library.

mass-translator About Input english text, then translate it between languages n times using the Deep Translator Python Library. How to Use Install dep

2 Mar 04, 2022

Study German declensions (dER nettE Mann, ein nettER Mann, mit dEM nettEN Mann, ohne dEN nettEN Mann ...) Generate as many exercises as you want using the incredible power of SPACY!

4 Jul 20, 2022

Bu Chatbot, Konya Bilim Merkezi Yen için tasarlanmış olan bir projedir.

chatbot Bu Chatbot, Konya Bilim Merkezi Yeni Ufuklar Sergisi için 2021 Yılında tasarlanmış olan bir projedir. Chatbot Python ortamında yazılmıştır. Sö

1 Feb 23, 2022

Beautiful visualizations of how language differs among document types.

Scattertext 0.1.0.0 A tool for finding distinguishing terms in corpora and displaying them in an interactive HTML scatter plot. Points corresponding t

2k Dec 27, 2022

The model is designed to train a single and large neural network in order to predict correct translation by reading the given sentence.

Neural Machine Translation communication system The model is basically direct to convert one source language to another targeted language using encode

7 Sep 22, 2022

This repository contains the code for running the character-level Sandwich Transformers from our ACL 2020 paper on Improving Transformer Models by Reordering their Sublayers.

Improving Transformer Models by Reordering their Sublayers This repository contains the code for running the character-level Sandwich Transformers fro

53 Sep 26, 2022

This repository has a implementations of data augmentation for NLP for Japanese.

Related tags

Overview

daaja

Install

How to use

Command

In Python

Command

In Python

Comments

Release 1.2.0

Release 1.1.1

Releases(v0.0.7)

v0.0.7(Oct 24, 2022)

Changes

v0.0.6(Mar 3, 2022)

Changes

📖 Documentation

v0.0.5(Feb 27, 2022)

Changes

💪 Enhancement

📖 Documentation

v0.0.4(Feb 21, 2022)

Changes

💪 Enhancement

📖 Documentation

v0.0.3(Feb 13, 2022)

v0.0.2(Feb 13, 2022)

Owner

Koga Kobayashi

Enterprise Scale NLP with Hugging Face & SageMaker Workshop series

File-based TF-IDF: Calculates keywords in a document, using a word corpus.

LeBenchmark: a reproducible framework for assessing SSL from speech

A list of NLP(Natural Language Processing) tutorials

Open-source offline translation library written in Python. Uses OpenNMT for translations

ChessCoach is a neural network-based chess engine capable of natural-language commentary.

justCTF [*] 2020 challenges sources

Input english text, then translate it between languages n times using the Deep Translator Python Library.

Study German declensions (dER nettE Mann, ein nettER Mann, mit dEM nettEN Mann, ohne dEN nettEN Mann ...) Generate as many exercises as you want using the incredible power of SPACY!

Bu Chatbot, Konya Bilim Merkezi Yen için tasarlanmış olan bir projedir.

Beautiful visualizations of how language differs among document types.

The model is designed to train a single and large neural network in order to predict correct translation by reading the given sentence.

This repository contains the code for running the character-level Sandwich Transformers from our ACL 2020 paper on Improving Transformer Models by Reordering their Sublayers.

gaiic2021-track3-小布助手对话短文本语义匹配复赛rank3、决赛rank4

Toward a Visual Concept Vocabulary for GAN Latent Space, ICCV 2021

Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents

Concept Modeling: Topic Modeling on Images and Text

Sequence modeling benchmarks and temporal convolutional networks

Easy Language Model Pretraining leveraging Huggingface's Transformers and Datasets

Based on 125GB of data leaked from Twitch, you can see their monthly revenues from 2019-2021