Skip to content
84

awesome-japanese-nlp-resources

A curated list of resources for Japanese natural language processing (NLP): Python libraries, LLMs, dictionaries, corpora, and datasets. Includes Claude Code and Codex skills to search resources.

1k stars53 forks901 entriesLast push Sep 28, 2026 (yesterday)License CC0-1.0

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

The latest additions

jev-auto-ime

打っている言葉が日本語か英語かをJevに聞いて、Macの入力モードを切り替える道具。個人利用のみ。

jpnorm

日本語テキスト正規化ライブラリ (Rust core + Python)。neologdn 互換・用途別プリセット・URL 保護・カスタム辞書 / Fast, configurable Japanese text normalization

adlib

ADLIB: Japanese ASR benchmark framework with language-aware evaluationb

Python library >Morphology analysis

sudachi.rs

SudachiPy 0.6* and above are developed as Sudachi.rs.

Janome

Japanese morphological analysis engine written in pure Python

mecab-python3

mecab-python. you can find original version here:http://taku910.github.io/mecab/

mecab

This repository is for building Windows 64-bit MeCab binary and improving MeCab Python binding.

fugashi

A Cython MeCab wrapper for fast, pythonic Japanese tokenization and morphological analysis.

nagisa

A Japanese tokenizer based on recurrent neural networks

pyknp

A Python Module for JUMAN++/KNP

Mykytea-python

Python wrapper for KyTea

konoha

Konoha: Simple wrapper of Japanese Tokenizers

natto-py

natto-py combines the Python programming language with MeCab, the part-of-speech and morphological analyzer for the Japanese language.

rakutenma-python

Rakuten MA (Python version)

python-vaporetto

Vaporetto is a fast and lightweight pointwise prediction based tokenizer. This is a Python wrapper for Vaporetto.

dango

An easy to use tokenizer for Japanese text, aimed at language learners and non-linguists

rhoknp

Yet another Python binding for Juman++/KNP

python-vibrato

Viterbi-based accelerated tokenizer (Python wrapper)

jagger-python

Python binding for Jagger(C++ implementation of Pattern-based Japanese Morphological Analyzer)

Mecari

Mecari (Japanese Morphological Analysis with Graph Neural Networks)

SudachiPy

🔴 october 2022

Python library >Parsing

ginza

A Japanese NLP Library using spaCy as framework based on Universal Dependencies

cabocha

Yet Another Japanese Dependency Structure Analyzer

UniDic2UD

Tokenizer POS-tagger Lemmatizer and Dependency-parser for modern and contemporary Japanese

camphr

Camphr - NLP libary for creating pipeline components

SuPar-UniDic

Tokenizer POS-tagger Lemmatizer and Dependency-parser for modern and contemporary Japanese with BERT models

depccg

A* CCG Parser with a Supertag and Dependency Factored Model

bertknp

A Japanese dependency parser based on BERT

esupar

Tokenizer POS-Tagger and Dependency-parser with BERT/RoBERTa/DeBERTa models for Japanese and other languages

yomikata

Heteronym disambiguation library using a fine-tuned BERT model.

jdepp-python

Python binding for J.DepP(C++ implementation of Japanese Dependency Parsers)

lightblue

A CCG parser for Japanese with DTS-representations

natsume-simple

natsume-simpleは日本語の係り受け関係検索システム

jdeppy

Python wrapper for J.DepP, fast Japanese Dependency Parser

Python library >Converter

pykakasi

Lightweight converter from Japanese Kana-kanji sentences into Kana-Roman.

cutlet

Japanese to romaji converter in Python

alphabet2kana

Convert English alphabet to Katakana

Convert-Numbers-to-Japanese

Converts Arabic numerals, or 'western' style numbers, to a Japanese context.

mozcpy

Mozc for Python: Kana-Kanji converter

jamorasep

Japanese text parser to separate Hiragana/Katakana string into morae (syllables).

text2phoneme

日本語文を音素列へ変換するスクリプト

jntajis-python

A fast character conversion and transliteration library based on the scheme defined for Japan National Tax Agency (国税庁) 's

wiredify

Convert japanese kana from ba-bi-bu-be-bo into va-vi-vu-ve-vo

mecab-text-cleaner

Simple Python package (CLI/Python API) for getting japanese readings (yomigana) and accents using MeCab.

pynormalizenumexp

数量表現や時間表現の抽出・正規化を行うNormalizeNumexpのPython実装

Jusho

Easy wrapper for the postal code data of Japan

yurenizer

Japanese text normalizer that resolves spelling inconsistencies. (日本語表記揺れ解消ツール)

e2k

A tool for automatic English to Katakana conversion

alkana.py

A tool to get the katakana reading of an alphabetical string.

englishtokanaconverter

英語文字列をカタカナに変換するプログラム

kanjiconv

Kanji Converter to Hiragana, Katakana, Roman alphabet.

kanjize

Kanjize(カンジャイズ): Easy converter between Kanji-Number and Integer

Python library >Preprocessor

neologdn

Japanese text normalizer for mecab-neologd

jaconv

Pure-Python Japanese character interconverter for Hiragana, Katakana, Hankaku, and Zenkaku

mojimoji

A fast converter between Japanese hankaku and zenkaku characters

text-cleaning

A powerful text cleaner for Japanese web texts

HojiChar

複数の前処理を構成して管理するテキスト前処理ツール

utsuho

Utsuho is a Python module that facilitates bidirectional conversion between half-width katakana and full-width katakana in Japanese.

python-habachen

Yet Another Fast Japanese String Converter

kairyou

Quickly preprocesses Japanese text using NLP/NER from SpaCy for Japanese translation or other NLP tasks.

Python library >Sentence splitter

Bunkai

Sentence boundary disambiguation tool for Japanese texts (日本語文境界判定器)

japanese-sentence-breaker

Japanese Sentence Breaker

sengiri

Yet another sentence-level tokenizer for the Japanese text

budoux

Standalone. Small. Language-neutral. BudouX is the successor to Budou, the machine learning powered line break organizer tool.

ja_sentence_segmenter

japanese sentence segmentation library for python

hasami

A tool to perform sentence segmentation on Japanese text

kuzukiri

Japanese Text Segmenter for Python written in Rust

ja-senter-benchmark

Comparison of Japanese Sentence Segmentation Tools

fast-bunkai

Japanese sentence splitting(日本語文境界判定器), 40–250× faster via a Rust-accelerated Python library with near-perfect API compatibility with megagonlabs/bunkai.

Python library >Sentiment analysis

oseti

Dictionary based Sentiment Analysis for Japanese

negapoji

Japanese negative positive classification.日本語文書のネガポジを判定。

pymlask

Emotion analyzer for Japanese text

asari

Japanese sentiment analyzer implemented in Python.

kotobacore

Japanese semantic understanding engine for LLM, RAG, SNS analysis, and AI agents ? dependency-free tokenizer + Plutchik emotion + intent + RAG keywords.

Python library >Machine translation

jparacrawl-finetune

An example usage of JParaCrawl pre-trained Neural Machine Translation (NMT) models.

JASS

JASS: Japanese-specific Sequence to Sequence Pre-training for Neural Machine Translation (LREC2020) & Linguistically Driven Multi-Task Pre-Training for Low-Resource Neural Machine Translation (ACM TALLIP)

PheMT

A phenomenon-wise evaluation dataset for Japanese-English machine translation robustness. The dataset is based on the MTNT dataset, with additional annotations of four linguistic phenomena; Proper Noun, Abbreviated Noun, Colloquial Expression, and Variant. COLING 2020.

VISA

An ambiguous subtitles dataset for visual scene-aware machine translation

plamo-translate-cli

A command-line interface for translation using the plamo-2-translate model with local execution.

Python library >Named entity recognition

namaco

Character Based Named Entity Recognition.

entitypedia

Entitypedia is an Extended Named Entity Dictionary from Wikipedia.

noyaki

Converts character span label information to tokenized text-based label information.

bert-japanese-ner-finetuning

Code to perform finetuning of the BERT model. BERTモデルのファインチューニングで固有表現抽出用タスクのモデルを作成・使用するサンプルです

joint-information-extraction-hs

詳細なアノテーション基準に基づく症例報告コーパスからの固有表現及び関係の抽出精度の推論を行うコード

pygeonlp

pygeonlp, A python module for geotagging Japanese texts.

bert-ner-japanese

BERTによる日本語固有表現抽出のファインチューニング用プログラム

huggingface-finetune-japanese

Examples to finetune encoder-only and encoder-decoder transformers for Japanese language (Hugging Face) Resources

novelanalysisbyner

BERTのfine-tuningによる固有表現抽出

Python library >OCR

Manga OCR

About Optical character recognition for Japanese text, with the main focus being Japanese manga

mokuro

Read Japanese manga inside browser with selectable text.

handwritten-japanese-ocr

Handwritten Japanese OCR demo using touch panel to draw the input text using Intel OpenVINO toolkit

OCR_Japanease

日本語OCR

ndlocr_cli

NDLOCRのアプリケーション

donut

Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022

JMTrans

manga translator - get japanese manga from url to translate manga image

Kindai-OCR

OCR system for recognizing modern Japanese magazines

text_recognition

NDLOCR用テキスト認識モジュール

Poricom

Optical character recognition in manga images. Manga OCR desktop application

owocr

Optical character recognition for Japanese text

yomitoku

Yomitoku is an AI-powered document image analysis package designed specifically for the Japanese language.

findtextcenternet

Japanese OCR with CenterNet

simple-ocr-for-manga

A simple OCR for manga (Japanese traditional and Japanese vertical)

jp-ocr-evaluation

日本語の文章画像に対するOCRの性能を評価

paddleocr-vl-sft-for-japanese-manga-on-rtx-3060

Fine-tune PaddleOCR-VL on the Manga109s dataset for Japanese manga text recognition. The base model struggles with vertical Japanese text reading order in manga. After fine-tuning, the model correctly handles manga-specific text layouts.

MangaOCR

A lightweight OCR model for Japanese text, especially in Manga

meikiocr

high-speed, high-accuracy, local ocr for japanese video games

meikipop

universal japanese ocr popup dictionary for windows, linux and macos

In 2 lists

Python library >Tool for pretrained models

JGLUE

JGLUE: Japanese General Language Understanding Evaluation

ginza-transformers

Use custom tokenizers in spacy-transformers

t5_japanese_dialogue_generation

T5による会話生成

japanese_text_classification

To investigate various DNN text classifiers including MLP, CNN, RNN, BERT approaches.

Japanese-BERT-Sentiment-Analyzer

Deploying sentiment analysis server with FastAPI and BERT

jmlm_scoring

Masked Language Model-based Scoring for Japanese and Vietnamese

allennlp-shiba-model

AllenNLP integration for Shiba: Japanese CANINE model

evaluate_japanese_w2v

script to evaluate pre-trained Japanese word2vec model on Japanese similarity dataset

gector-ja

BERT-based GEC tagging for Japanese

Japanese-BPEEncoder

Japanese-BPEEncoder

Japanese-BPEEncoder_V2

Japanese-BPEEncoder Version 2

transformer-copy

日本語文法誤り訂正ツール

nagisa_bert

A BERT model for nagisa

JGLUE-benchmark

Training and evaluation scripts for JGLUE, a Japanese language understanding benchmark

jptranstokenizer

Japanese Tokenizer for transformers library

jp-stable

JP Language Model Evaluation Harness

compare-ja-tokenizer

How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in Japanese-ACL SRW 2023

lm-evaluation-harness-jp-stable

A framework for few-shot evaluation of autoregressive language models.

llm-lora-classification

llm-lora-classification

rinna_gpt-neox_ggml-lora

The repository contains scripts and merge scripts that have been modified to adapt an Alpaca-Lora adapter for LoRA tuning when assuming the use of the "rinna/japanese-gpt-neox..." [gpt-neox] model converted to ggml.

In 2 lists

japanese-llm-roleplay-benchmark

このリポジトリは日本語LLMのキャラクターロールプレイに関する性能を評価するために作成しました。

japanese-llm-ranking

This repository supports YuzuAI's Rakuda leaderboard of Japanese LLMs, which is a Japanese-focused analogue of LMSYS' Vicuna eval.

llm-jp-eval

このツールは、複数のデータセットを横断して日本語の大規模言語モデルを自動評価するものです.

llm-jp-sft

This repository contains the code for supervised fine-tuning of LLM-jp models.

llm-jp-tokenizer

LLM勉強会(LLM-jp)で開発しているLLM用のトークナイザー関連をまとめたリポジトリです.

japanese-lm-fin-harness

Japanese Language Model Financial Evaluation Harness

ja-vicuna-qa-benchmark

Japanese Vicuna QA Benchmark

swallow-evaluation

Swallowプロジェクト 大規模言語モデル 評価スクリプト

swallow-evaluation-instruct

Swallowプロジェクト 事後学習ずみ大規模言語モデル 評価フレームワーク

pretrained_doc2vec_ja

pretrained doc2vec models on Japanese Wikipedia

pl-bert-ja

A repository of Japanese Phoneme-Level BERT

Show 211 items

namedivider-python

A tool for dividing the Japanese full name into a family name and a given name.

asa-python

A curated list of resources dedicated to Python libraries of NLP for Japanese

python_asa

python版日本語意味役割付与システム(ASA)

toiro

A comparison tool of Japanese tokenizers

ja-timex

自然言語で書かれた時間情報表現を抽出/規格化するルールベースの解析器

JapaneseTokenizers

A set of metrics for feature selection from text data

daaja

This repository has implementations of data augmentation for NLP for Japanese.

accel-brain-code

The purpose of this repository is to make prototypes as case study in the context of proof of concept(PoC) and research and development(R&D) that I have written in my website. The main research topics are Auto-Encoders in relation to the representation learning, the statistical machine learning…

kyoto-reader

A processor for KyotoCorpus, KWDLC, and AnnotatedFKCCorpus

nlplot

Visualization Module for Natural Language Processing

rake-ja

Rapid Automatic Keyword Extraction algorithm for Japanese

jel

Japanese Entity Linker.

MedNER-J

Latest version of MedEX/J (Japanese disease name extractor)

zunda-python

Zunda: Japanese Enhanced Modality Analyzer client for Python.

AIO2_DPR_baseline

https://www.nlp.ecei.tohoku.ac.jp/projects/aio/

showcase

A PyTorch implementation of the Japanese Predicate-Argument Structure (PAS) analyser presented in the paper of Matsubayashi & Inui (2018) with some improvements.

darts-clone-python

Darts-clone python binding

jrte-corpus_example

Example codes for Japanese Realistic Textual Entailment Corpus

desuwa

Feature annotator to morphemes and phrases based on KNP rule files (pure-Python)

HotPepperGourmetDialogue

Restaurant Search System through Dialogue in Japanese.

nlp-recipes-ja

Samples codes for natural language processing in Japanese

Japanese_nlp_scripts

Small example scripts for working with Japanese texts in Python

DNorm-J

Japanese version of DNorm

pyknp-eventgraph

EventGraph is a development platform for high-level NLP applications in Japanese.

ishi

Ishi: A volition classifier for Japanese

python-npylm

ベイズ階層言語モデルによる教師なし形態素解析

python-npycrf

条件付確率場とベイズ階層言語モデルの統合による半教師あり形態素解析

unsupervised-pos-tagging

教師なし品詞タグ推定

negima

Negima is a Python package to extract phrases in Japanese text by using the part-of-speeches based rules you defined.

YouyakuMan

Extractive summarizer using BertSum as summarization model

japanese-numbers-python

A parser for Japanese number (Kanji, arabic) in the natural language.

kantan

Lookup japanese words by radical patterns

make-meidai-dialogue

Get Japanese dialogue corpus

japanese_summarizer

A summarizer for Japanese articles.

chirptext

ChirpText is a collection of text processing tools for Python.

yubin

Japanese Address Munger

jawiki-cleaner

Japanese Wikipedia Cleaner

japanese2phoneme

A python library to convert Japanese to phoneme.

anlp_nlp2021_d3-1

This repository contains codes related to the experiments in "An Experimental Evaluation of Japanese Tokenizers for Sentiment-Based Text Classification"

aozora_classification

This project aims to classify Japanese sentence to how well similar to some Japanese classical writers, such as Soseki Natsume, Ogai Mori, Ryunosuke Akutagawa and so on.

aozora-corpus-generator

Generates plain or tokenized text files from the Aozora Bunko

JLM

A fast LSTM Language Model for large vocabulary language like Japanese and Chinese

NTM

Testing of Neural Topic Modeling for Japanese articles

EN-JP-ML-Lexicon

This is a English-Japanese lexicon for Machine Learning and Deep Learning terminology.

text-generation

Easy-to-use scripts to fine-tune GPT-2-JA with your own texts, to generate sentences, and to tweet them automatically.

chainer_nic

Neural Image Caption (NIC) on chainer, its pretrained models on English and Japanese image caption datasets.

unihan-lm

The official repository for "UnihanLM: Coarse-to-Fine Chinese-Japanese Language Model Pretraining with the Unihan Database", AACL-IJCNLP 2020

mbart-finetuning

Code to perform finetuning of the mBART model.

xvector_jtubespeech

xvector model on jtubespeech

TinySegmenterMaker

TinySegmenter用の学習モデルを自作するためのツール.

Grongish

日本語とグロンギ語の相互変換スクリプト

WordCloud-Japanese

WordCloudでの日本語文章をMecab(形態素解析エンジン)を使用せずに形態素解析チックな表示を実現するスクリプト

snark

日本語ワードネットを利用したDBアクセスライブラリ

toEmoji

日本語文を絵文字だけの文に変換するなにか

termextract

専門用語抽出アルゴリズムの実装の練習

JDT-with-KenLM-scoring

Japanese-Dialog-Transformerの応答候補に対して、KenLMによるN-gram言語モデルでスコアリングし、フィルタリング若しくはリランキングを行う。

mixture-of-unigram-model

Mixture of Unigram Model and Infinite Mixture of Unigram Model in Python. (混合ユニグラムモデルと無限混合ユニグラムモデル)

hidden-markov-model

Hidden Markov Model (HMM) and Infinite Hidden Markov Model (iHMM) in Python. (隠れマルコフモデルと無限隠れマルコフモデル)

Ngram-language-model

Ngram language model in Python. (Nグラム言語モデル)

ASRDeepSpeech

Automatic Speech Recognition with deepspeech2 model in pytorch with support from Zakuro AI.

neural_ime

Neural IME: Neural Input Method Engine

neural_japanese_transliterator

Can neural networks transliterate Romaji into Japanese correctly?

tinysegmenter

tokenizer specified for Japanese

AugLy-jp

Data Augmentation for Japanese Text on AugLy

furigana4epub

A Python script for adding furigana to Japanese epub books using Mecab and Unidic.

PyKatsuyou

Japanese verb/adjective inflections tool

jageocoder

Pure Python Japanese address geocoder

nksnd

New kana-kanji conversion engine

JaMIE

A Japanese Medical Information Extraction Toolkit

fasttext-vs-word2vec-on-twitter-data

fasttextとword2vecの比較と、実行スクリプト、学習スクリプトです

minimal-search-engine

最小のサーチエンジン/PageRank/tf-idf

5ch-analysis

5chの過去ログをスクレイピングして、過去流行った単語(ex, 香具師, orz)などを追跡調査

tweet_extructor

Twitter日本語評判分析データセットのためのツイートダウンローダ

japanese-word-aggregation

Aggregating Japanese words based on Juman++ and ConceptNet5.5

jinf

A Japanese inflection converter

kwja

A unified language analyzer for Japanese

mlm-scoring-transformers

Reproduced package based on Masked Language Model Scoring (ACL2020).

ClipCap-for-Japanese

[PyTorch] ClipCap for Japanese

SAT-for-Japanese

[PyTorch] Show, Attend and Tell for Japanese

cihai

Python library for CJK (Chinese, Japanese, and Korean) language dictionary

marine

MARINE : Multi-task leaRnIng-based JapaNese accent Estimation

whisper-asr-finetune

Finetuning Whisper ASR model

radicalchar

部首文字正規化ライブラリ

akaza

Yet another Japanese IME for IBus/Linux

posuto

Japanese postal code data.

tacotron2-japanese

Tacotron2 implementation of Japanese

ibus-hiragana

ひらがなIME for IBus

furiganapad

ふりがなパッド

chikkarpy

Japanese synonym library

ja-tokenizer-docker-py

Mecab + NEologd + Docker + Python3

JapaneseEmbeddingEval

JapaneseEmbeddingEval

shuwa

Extend GNOME On-Screen Keyboard for Input Methods

japanese-nli-model

This repository provides the code for Japanese NLI model, a fine-tuned masked language model.

tra-fugu

A tool for Japanese-English translation and English-Japanese translation by using FuguMT

fugumt

ぷるーふおぶこんせぷと で公開した機械翻訳エンジンを利用する翻訳環境です。 フォームに入力された文字列の翻訳、PDFの翻訳が可能です。

JaSPICE

JaSPICE: Automatic Evaluation Metric Using Predicate-Argument Structures for Image Captioning Models

Retrieval-based-Voice-Conversion-WebUI-JP-localization

jp-localization

pyopenjtalk

Python wrapper for OpenJTalk

yomigana-ebook

Make learning Japanese easier by adding readings for every kanji in the eBook

N46Whisper

Whisper based Japanese subtitle generator

japanese_llm_simple_webui

Rinna-3.6B、OpenCALM等の日本語対応LLM(大規模言語モデル)用の簡易Webインタフェースです

pdf-translator

pdf-translator translates English PDF files into Japanese, preserving the original layout.

japanese_qa_demo_with_haystack_and_es

Haystack + Elasticsearch + wikipedia(ja) を用いた、日本語の質問応答システムのサンプル

mozc-devices

Automatically exported from code.google.com/p/mozc-morse

vits-japros-webui

日本語TTS(VITS)の学習と音声合成のGradio WebUI

ja-law-parser

A Japanese law parser

dictation-kit

Japanese dictation kit using Julius

julius4seg

Juliusを使ったセグメンテーション支援ツール

voicevox_engine

無料で使える中品質なテキスト読み上げソフトウェア、VOICEVOXの音声合成エンジン

LLaVA-JP

LLaVA-JP is a Japanese VLM trained by LLaVA method

RAG-Japanese

Open source RAG with Llama Index for Japanese LLM in low resource settting

bertjsc

Japanese Spelling Error Corrector using BERT(Masked-Language Model). BERTに基づいて日本語校正

llm-leaderboard

Project of llm evaluation to Japanese tasks

jglue-evaluation-scripts

Training and evaluation scripts for JGLUE, a Japanese language understanding benchmark

BLIP2-Japanese

Modifying LAVIS' BLIP2 Q-former with models pretrained on Japanese datasets.

wikipedia-passages-jawiki-embeddings-utils

wikipedia 日本語の文を、各種日本語の embeddings や faiss index へと変換するスクリプト等。

simple-simcse-ja

Exploring Japanese SimCSE

gpt4-autoeval

GPT-4 を用いて、言語モデルの応答を自動評価するスクリプト

t5-japanese

日本語T5モデル

japanese_llm_eval

A repo for evaluating Japanese LLMs ・ 日本語LLMを評価するレポ

jmteb

The evaluation scripts of JMTEB (Japanese Massive Text Embedding Benchmark)

pydomino

日本語音声に対して音素ラベルをアラインメントするためのツールです

easynovelassistant

軽量で規制も検閲もない日本語ローカル LLM『LightChatAssistant-TypeB』による、簡単なノベル生成アシスタントです。ローカル特権の永続生成 Generate forever で、当たりガチャを積み上げます。読み上げにも対応。

clip-japanese

日本語データセットでのqlora instruction tuning学習サンプルコード

rime-jaroomaji

Japanese rōmaji input schema for Rime IME

deep-question-generation

深層学習を用いたクイズ自動生成(日本語T5モデル)

magpie-nemotron

Magpieという手法とNemotron-4-340B-Instructを用いて合成対話データセットを作るコード

qlora_ja

日本語データセットでのqlora instruction tuning学習サンプルコード

mozcdic-ut-jawiki

Mozc UT Jawiki Dictionary is a dictionary generated from the Japanese Wikipedia for Mozc.

shisa-v2

Japanese / English Bilingual LLM

llm-translator

Mixtral-based Ja-En (En-Ja) Translation model

llm-jp-asr

Whisperのデコーダをllm-jp-1.3b-v1.0に置き換えた音声認識モデルを学習させるためのコード

rag-japanese

Open source RAG with Llama Index for Japanese LLM in low resource settting

monaka

A Japanese Parser (including historical Japanese)

jp-translate.cloud

A state-of-the-art open-source Japanese <--> English machine translation system based on the latest NMT research.

substring-word-finder

連続部分文字列の単語判定を行います

heron-vlm-leaderboard

This project is a benchmarking tool for evaluating and comparing the performance of various Vision Language Models (VLMs). It uses two datasets: LLaVA-Bench-In-the-Wild and Japanese HERON Bench to measure model performance.

text2dataset

Easily turn large English text datasets into Japanese text datasets using open LLMs.

mecab-web-api

MeCabを利用した日本語形態素解析WebAPI

mecab_controller

Mecab wrapper to generate furigana readings.

vits

VITSによるテキスト読み上げ器&ボイスチェンジャー

akari_chatgpt_bot

音声認識、文章生成、音声合成を使って対話するチャットボットアプリ

kudasai

Streamlining Japanese-English Translation with Advanced Preprocessing and Integrated Translation Technologies

mecab-visualizer

MeCabの形態素解析結果を可視化するツール

add-dictionary

OpenJTalkのユーザ辞書をGUIで追加するアプリ

j-moshi

J-Moshi: A Japanese Full-duplex Spoken Dialogue System

jatts

JATTS: Japanese TTS (for research)

tsukasa-speech

a Frontier Japanese Speech Generation net

symptom-expression-search

ElasticsearchやGiNZA、患者表現辞書を使った患者表現揺れ吸収する意味構造検索を試した

llm-jp-judge

生成自動評価を行うためのPythonツール

asagi-vlm-colaboratory-sample

Colaboratory上でAsagi(合成データセットを活用した大規模日本語VLM)をお試しするサンプル

llm-jp-eval-mm

This tool automatically evaluates Japanese multi-modal large language models across multiple datasets.

manga109api

Simple python API to read annotation data of Manga109

fastrtc-jp

fastrtc用の日本語TTSとSTT追加キット

whisper-transcription

Pythonを使用したWhisperモデルによる音声文字起こしツール

pocket-researcher

LLMを活用した自律調査エージェント。手軽に情報収集、概要把握。

jtransbench

A tool to easily benchmark Japanese translation skills

easyllasa

EasyLlasa は 5~15秒の日本語音声と日本語テキストから日本語音声を生成する TSTS (TextSpeechToSpeech) です。

kanjikana-model

氏名漢字カナ突合モデル

deep-openreview-research-ja

OpenReview論文を自動で発見・分析する日本語対応AIエージェント

pitchbench

Experimental Japanese pitch accent based LLM Benchmark

mini-transformer-from-scratch

English to Japanese Transformer from scratch

vv_core_inference

VOICEVOXのコア内で用いられているディープラーニングモデルの推論コード

pyopenjtalk-plus

pyopenjtalk-plus: A Python wrapper for OpenJTalk with additional improvements

japanese_spelling_correction

Japanese Spelling Correction

py-kaomoji

python kaomoji

llm-jp-vila

This repository contains the code for training llm-jp/llm-jp-3-vila-14b, modified from VILA repository.

kanjivg-radical

kanjivg-radical

japanese-wordnet-visualization

This project visualizes the Japanese Wordnet (日本語ワードネット) with web application built by Django

piper-plus

Enhanced Piper TTS with Japanese support, WebAssembly, multi-GPU training, and quality improvements.

Japanera

Easy Tools for Japanese Era System

bert-abstractive-text-summarization

Japanese Sentence Summarization with BERT

kyujipy

A Python library to convert Japanese texts from Shinjitai (新字体) to Kyujitai (舊字體) and vice versa

jitenbot

Web crawler for creating personal copies of Japanese dictionaries

ja-icd10

ICD-10 国際疾病分類の日本語情報を扱うためのPythonパッケージ

pl-bert-vits2

VITS2 using Phoneme-Level Japanese BERT

ndc_predictor

NDCPredictorの機械学習モデル(書誌情報から日本十進分類を推測するfastTextの学習済みモデル)

pfmt-bench-fin-ja

pfmt-bench-fin-ja: Preferred Multi-turn Benchmark for Finance in Japanese

marine-plus

MARINE : Multi-task leaRnIng-based JapaNese accent Estimation (Also supported Windows)

ja-tokenizer-benchmark

Compare the speed of various Japanese tokenizers in Python.

yat

yat: Yet Another Tokenizer for Japanese NLP

igakuqa119

Evaluating LLMs on the 119th Japanese Medical Licensing Examination

japanese-luw-tokenizer

Japanese Long-Unit-Word Tokenizer with RemBertTokenizerFast of Transformers

ibus-jig

ibus-jig: Japanese-language Input-method using GPT-4

jp-stopword-filter

A lightweight Python library designed to filter stopwords from Japanese text based on customizable rules.

yasumail

Synthetic Japanese business email generator for ML training data

himotoki

A Python-based Japanese Tokenizer, Dictionary, Morphological Analyzer and Romanization Tool. Based on JMDict for Language Learning.

diafill-toolkit

A toolkit for synthesizing filler-rich, short-utterance Japanese dialogue scripts for speech-based interaction using Large Language Models (LLMs) This project is designed to generate data in two phases: Seed Generation (metadata creation) and Dialogue Generation (script creation).

eval_vertical_ja

Evaluating Multimodal Large Language Models on Vertically Written Japanese Text

jp-llm-corpus-pii-filter

本コードは,大規模言語モデル(LLM)の学習用コーパスから,個人情報の中でも特に配慮が求められる「要配慮個人情報」をフィルタリングするためのものです.

Novel2DialCorpus

小説テキストから雑談対話コーパスを構築する手法

OneCompression

富士通研究所による LLM 向け後学習量子化 (PTQ) パイプライン。QEP (NeurIPS 2025)、ILP 混合精度、回転前処理、vLLM プラグインを統合。論文: arXiv:2603.28845。

In 2 lists

manga-translator

Translate text found within speech bubbles in manga images.

shirabe-address-api

Shirabe Address API — Japanese address normalization for AI agents (Cloudflare Workers + Fly.io NRT, abr-geocoder backed)

medical-paper-summarizer-public

毎日PubMedから最新の循環器内科論文を自動収集・AI要約してGmailに届けるシステム

Irodori-TTS

A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control

sarashina2.2-tts

Sarashina2.2-TTS is a Japanese-centric text-to-speech system built on a large language model, developed by SB Intuitions.

manga-translator

A manga translator created using Yolov8, manga_ocr and deep-translator.

jp-tl-bench

Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation

novel2hermes_jp

メモリ機能が強力なhermes-agentと、日本語検索に強い外部メモリvecmemoriを活かし、長文に耐える小説を企画/プロッティング/執筆するためのskills.md

moshi-finetune

Fine-tuning Moshi/J-Moshi on your own spoken dialogue data

simple-evals-mm

A lightweight library for evaluating vision language models on English and Japanese benchmarks.

medvoice-jp-asr

MedVoice-JP LoRA60: 日本語医療特化ASR (whisper-small + LoRA, 60話者) と国内初の医療ASRベンチマーク (フルFT版 FT60 も収録)

open-zeimu-mcp

OSS MCP server for Japanese tax primary sources (法令・通達・タックスアンサー・質疑応答事例・文書回答事例・裁決事例)

llm-jp-moshi

LLM-jp-Moshi: Japanese Full-duplex Spoken Dialogue Models

shirabe-sdk-python

Official Python SDK for the Shirabe Japan data APIs — ready-made LangChain / OpenAI Agents SDK tools for Japanese name reading (JMnedict-backed candidates), name splitting, address normalization, corporate number lookup, and calendar (rokuyo). Zero-dependency core.

fuseji

日本語特化のPII検出・マスキングミドルウェア(LLMオブザーバビリティ向け)

moine

Romanization-aware string comparison for Japanese and Mandarin Chinese.

bpe2regex

BPE tokenizer をクソデカ正規表現に変換する意味わからんやつ

sanoTTS-jp

559 K パラメータの日本語 TTS を ESP32-S3 で実時間合成。漢字かな交じり文の形態素解析・アクセント推定まで端末内で走る(M5Stack CoreS3 実機で確認)。推論は依存ゼロの C99、ブラウザ demo あり。arXiv:2608.21378 sanoTTS の日本語 clean-room 再実装。⚠️ コードは MIT ですが、配布モデルの重みは MIT ではありません(LICENSE-MODEL.md。出力に用途制限が伝播します)

jev-auto-ime

打っている言葉が日本語か英語かをJevに聞いて、Macの入力モードを切り替える道具。個人利用のみ。

C++ >Morphology analysis

mecab

Yet another Japanese morphological analyzer

jumanpp

Juman++ (a Morphological Analyzer Toolkit)

kytea

The Kyoto Text Analysis Toolkit for word segmentation and pronunciation estimation, etc.

juman

Japanese Morphological Analysis System JUMAN

C++ >Parsing

cabocha

Yet Another Japanese Dependency Structure Analyzer

knp

A Japanese Parser

Show 7 items

jsc

Joint source channel model for Japanese Kana Kanji conversion, Chinese pinyin input and CJE mixed input.

aquaskk

An input method without morphological analysis.

mozc

Mozc - a Japanese Input Method Editor designed for multi-platform

trimatch

Trimatch: An (Exact|Prefix|Approximate) String Matching Library

resembla

Resembla: Word-based Japanese similar sentence search library

corvusskk

▽▼ SKK-like Japanese Input Method Editor for Windows

mozuku

日本語文章の解析・校正を行う LSP サーバー。

Rust crate >Morphology analysis

lindera

A morphological analysis library.

vaporetto

Vaporetto: Very Accelerated POintwise pREdicTion based TOkenizer

goya

Japanese Morphological Analysis written in Rust

vibrato

vibrato: Viterbi-based accelerated tokenizer

yoin

A Japanese Morphological Analyzer written in pure Rust

mecab-rs

Safe Rust bindings for mecab a part-of-speech and morphological analyzer library

awabi

A morphological analyzer using mecab dictionary

kanpyo

Japanese Morphological Analyzer written in Rust

mecrab

A pure Rust implementation of a morphological analyzer compatible with MeCab dictionaries (IPADIC format).

Rust crate >Converter

wana_kana_rust

Utility library for checking and converting between Japanese characters - Hiragana, Katakana - and Romaji

unicode-jp-rs

A Rust library to convert Japanese Half-width-kana[半角カナ] and Wide-alphanumeric[全角英数] into normal ones

kana

[Mirror] CLI program for transliterating romaji text to either hiragana or katakana

kanaria

このライブラリは、ひらがな・カタカナ、半角・全角の相互変換や判別を始めとした機能を提供します。

japanese-address-parser

日本の住所を都道府県/市区町村/町名/その他に分割するライブラリです

yosina

Yosina is a transliteration library deals with the letters and symbols used in Japanese writing.

mojimoji-rs

Rust implementation of a fast converter between Japanese hankaku and zenkaku characters, mojimoji.

haqumei

A Japanese Grapheme-to-Phoneme (G2P) library.

ja-furigana

日本語フリガナ (ルビ) を扱う Rust 製ライブラリ + ローカル HTTP サーバー。ルールはすべてデータ駆動 (TOML)。

jpnorm

日本語テキスト正規化ライブラリ (Rust core + Python)。neologdn 互換・用途別プリセット・URL 保護・カスタム辞書 / Fast, configurable Japanese text normalization

Rust crate >Search engine library

lindera-tantivy

Lindera tokenizer for Tantivy.

tantivy-vibrato

A Tantivy tokenizer using Vibrato.

sqlite-vaporetto

SQLite FTS5 extension for fast Japanese full-text search with 🛥Vaporetto / Vaporetto による高速な日本語全文検索を SQLite FTS5 で実現する拡張機能

duckdb-vaporetto

DuckDB extension for Japanese full-text search with 🛥Vaporetto / Vaporetto による DuckDB + 日本語全文検索拡張機能

lindera-sqlite

Lindera for SQLite FTS5 extention

Show 26 items

daachorse

A fast implementation of the Aho-Corasick algorithm using the compact double-array data structure in Rust.

find-simdoc

Finding all pairs of similar documents time- and memory-efficiently

crawdad

Rust library of natural language dictionaries using character-wise double-array tries.

tokenizer-speed-bench

Comparison code of various tokenizers

stringmatch-bench

Here provides benchmark tools to compare the performance of data structures for string matching.

vime

Using Vim as an input method for X11 apps

voicevox_core

無料で使える中品質なテキスト読み上げソフトウェア、VOICEVOXのコア

akaza

Yet another Japanese IME for IBus/Linux

Jotoba

A free online, self-hostable, multilang Japanese dictionary.

dvorakjp-romantable

Google 日本語入力用DvorakJPローマ字テーブル / DvorakJP Roman Table for Google Japanese Input

niinii

Japanese glossator for assisted reading of text using Ichiran

cskk

SKK (Simple Kana Kanji henkan) library

japanki

Learn Japanese vocabs 🇯🇵 by doing quizzes on CLI!

jpreprocess

Japanese text preprocessor for Text-to-Speech applications (OpenJTalk rewrite in rust language)

listup_precedent

裁判例のデータ一覧を裁判所のホームページ(https://www.courts.go.jp/index.html) をスクレイピングして生成するソフトウェア

jisho

Jisho is a CLI tool & Rust library that provides a Japanese-English dictionary.

kanalizer

英単語から読みを推測するライブラリ。

koharu

Automated manga translation tool with LLM, written in Rust.

In 2 lists

yomine

A Japanese vocabulary mining tool designed to help language learners mine new words and expressions.

matsuba

lightweight japanese ime written in rust

hujiang_dictionary

日本語辞書 by Rust, support Telegram bot, AWS Lambda and Cloudflare Workers. Support LLM and search RAG.

mecab-dic-converter

MeCab 用のコンパイル済み辞書を読み取り、解析・再構成し、最終的に Vibrato や Lindera のようなMecab互換 tokenizer が使える辞書へ変換するための Rust crate です。

jp-deinflector

A high-performance Rust crate for deinflecting Japanese words using perfect hash tables

mikke

日本語 Markdown ノートのローカル検索 CLI 👀 — BM25 全文検索 (SQLite FTS5) + optional なローカル semantic/hybrid。単一バイナリ・外部 API 不使用、AI コーディングエージェント向け。

suiko

日本語文書の自然さと読みやすさを再現可能に診断するRust CLI / Deterministic diagnostics for natural and readable Japanese writing

daac-bpe

Fast incremental BPE tokenizer based on daachorse

JavaScript >Morphology analysis

kuromoji.js

JavaScript implementation of Japanese morphological analyzer

rakutenma

Rakuten MA - morphological analyzer (word segmentor + PoS Tagger) for Chinese and Japanese written purely in JavaScript.

node-mecab-ya

Yet another mecab wrapper for nodejs

juman-bin

a User-Extensible Morphological Analyzer for Japanese. 日本語形態素解析システム

node-mecab-async

Asynchronous japanese morphological analyser using MeCab.

JavaScript >Converter

kuroshiro

Japanese language library for converting Japanese sentence to Hiragana, Katakana or Romaji with furigana and okurigana modes supported.

In 2 lists

kuroshiro-analyzer-kuromoji

Kuromoji morphological analyzer for kuroshiro.

hepburn

Node.js module for converting Japanese Hiragana and Katakana script to, and from, Romaji using Hepburn romanisation

japanese-numerals-to-number

Converts Japanese Numerals into number

jslingua

Javascript libraries to process text: Arabic, Japanese, etc.

WanaKana

Javascript library for detecting and transliterating Hiragana <--> Katakana <--> Romaji

node-romaji-name

Normalize and fix common issues with Romaji-based Japanese names.

kyujitai.js

Utility collections for making Japanese text old-fashioned

normalize-japanese-addresses

オープンソースの住所正規化ライブラリ。

jaconv

日本語文字変換ライブラリ (javascript)

romaji-conv

Convert romaji into hiragana

japanese-addresses-v2

全国の住所データAPI

jptext-to-emoji

テキストの単語を絵文字に変換する

japanese.js

Util collection for Japanese text processing. Hiraganize, Katakanize, and Romanize.

genshijin

About genshijin 原始人 🗿| Claude Code / Codex等AIエージェント 向け超圧縮コミュニケーションスキル。caveman の日本語版をベースに、日本語特有の冗長表現に最適化。

Show 26 items

bangumi-data

Raw data for Japanese Anime

In 2 lists

yomichan

Japanese pop-up dictionary extension for Chrome and Firefox.

proofreading-tool

GUIで動作する文書校正ツール GUI tool for textlinting.

kanjigrid

A web-app displaying the 2200 kanji characters taught in James Heisig's "Remembering the Kanji", 6th edition.

japanese-toolkit

Monorepo for Kanji, Furigana, Japanese DB, and others

analyze-desumasu-dearu

文の敬体(ですます調)、常体(である調)を解析するJavaScriptライブラリ

hatsuon

Japanese pitch accent utils

sentiment_ja_js

Sentiment Analysis in Japanese. sentiment_ja with JavaScript

mecab-ipadic-seed

mecab-ipadic seed dictionary reader

oskim

Extend GNOME On-Screen Keyboard for Input Methods

tweetMapping

東日本大震災発生から24時間以内につぶやかれたジオタグ付きツイートのデジタルアーカイブです。

pitch-accent

Predict pitch accent in Japanese

kana2ipa

「ひらがな」または「カタカナ」を日本語で発音する際の音声記号(IPA)に変換するコマンド

voicevox

無料で使える中品質なテキスト読み上げソフトウェア、VOICEVOXのエディター

In 2 lists

kamiya-codec

Towards a Japanese verb conjugator and deconjugator based on Taeko Kamiya's The Handbook of Japanese Verbs and The Handbook of Japanese Adjectives and Adverbs opuses.

closewords

最も似た単語を単語群から検索する日本語(漢字含む)対応のライブラリ

japanese-analyzer

Japanese Sentence Analyzer (日本語文章解析器)

japanese-furigana-normalize

Normalize Japanese Furigana

yama

acquire Japanese vocabulary on any website

kaitai

An application for analyzing Japanese sentence structure using AI. This tool visualizes how words and phrases relate to each other, showing grammatical relationships with interactive diagrams.

tsukeru-furigana-converter

Browser extension (Chrome/Edge/Firefox) that injects furigana into Japanese webpages on-demand; includes dictionary tooltips, JLPT filtering, and vocab/Anki export.

sudachi-synonyms-dictionary

Sudachi's synonyms dictionary

qmd-ja

Japanese-enhanced fork of qmd — Vaporetto WASM morphological tokenizer for accurate Japanese BM25 search

shirabe-sdk

Official TypeScript SDK for the Shirabe Japan data APIs — ready-made Vercel AI SDK / LangChain tools for Japanese name splitting/reading, address normalization, corporate number lookup, and calendar (rokuyo). Zero runtime dependencies in the core.

pii-ja-ner-onnx-demo

PII-JA NER Browser Demo

jev-semgrep

grep by meaning, across languages. TypeSafe Jev scores every line against a meaning; combine meanings with AND/OR/NOT. 意味で探す grep。日本語で英語を、英語で日本語を検索できる

In 2 lists

Go >Morphology analysis

kagome

Self-contained Japanese Morphological Analyzer written in pure Go

In 4 listsDetails

Show 9 items

ojosama

テキストを壱百満天原サロメお嬢様風の口調に変換します

nihongo

Japanese Dictionary

yomichan-import

External dictionary importer for Yomichan.

imas-ime-dic

THE IDOLM@STER words dictionary for Japanese IME (by imas-db.jp)

go-kakasi

Kanji transliteration to hiragana/katakana/romaji, in Go

go-moji

A Go library for Zenkaku/Hankaku conversion

ojichat

おじさんがLINEやメールで送ってきそうな文を生成する

name

Name Searcher in Japanese

jp-pii-detector

日本語個人情報検出器

Java >Morphology analysis

kuromoji

Kuromoji is a self-contained and very easy to use Japanese morphological analyzer designed for search

Sudachi

A Japanese Tokenizer for Business

SudachiDict

A lexicon for Sudachi

meval

形態素解析器性能評価システム MevAL

Show 9 items

kanjitomo-ocr

Java library for identifying Japanese characters from images

jakaroma

Java library and command-line tool to transliterate Japanese kanji to romaji (Latin alphabet)

kakasi-java

Kanji transliteration to hiragana/katakana/romaji, in Java

Kamite

A desktop language immersion companion for learners of Japanese

react-native-japanese-tokenizer

Async Japanese Tokenizer Native Plugin for React Native for iOS and Android

elasticsearch-analysis-japanese

Japanese analyzer uses kuromoji japanese tokenizer for ElasticSearch

moji4j

A Java library to converts between Japanese Hiragana, Katakana, and Romaji scripts.

neologdn-java

Japanese text normalizer for mecab-neologd

elasticsearch-sudachi

The Japanese analysis plugin for elasticsearch

Pretrained model >Word2Vec

japanese-words-to-vectors

Word2vec (word to vectors) approach for Japanese language using Gensim and Mecab.

chiVe

Japanese word embedding with Sudachi and NWJC

elmo-japanese

elmo-japanese

embedrank

Python Implementation of EmbedRank

aovec

Easy aozorabunko Word2Vec Builder - 青空文庫全書籍のWord2Vecビルダー+構築済みモデル

dependency-based-japanese-word-embeddings

This is a repository for the AI LAB article "係り受けに基づく日本語単語埋込 (Dependency-based Japanese Word Embeddings)" ( Article URL https://ai-lab.lapras.com/nlp/japanese-word-embedding/)

jawikivec

Yet Another Japanese-Wikipedia Entity Vectors

jawiki_word_vector_updater

最新の日本語Wikipediaのダンプデータから,MeCabを用いてIPA辞書と最新のNeologd辞書の両方で形態素解析を実施し,その結果に基づいた word2vec,fastText,GloVeの単語分散表現を学習するためのスクリプト

Pretrained model >Transformer based models

bert-japanese

BERT models for Japanese text.

bert-japanese

BERT with SentencePiece for Japanese text.

In 2 lists

SudachiTra

Japanese tokenizer for Transformers

japanese-dialog-transformers

Code for evaluating Japanese pretrained models provided by NTT Ltd.

shiba

Pytorch implementation and pre-trained Japanese model for CANINE, the efficient character-level transformer.

Dialog

A PyTorch Implementation of japanese chatbot using BERT and Transformer's decoder

language-pretraining

BERT and ELECTRA models of PyTorch implementations for Japanese text.

medbertjp

Trials of pre-trained BERT models for the medical domain in Japanese.

ILYS-aoba-chatbot

ILYS-aoba-chatbot

t5-japanese

Codes to pre-train Japanese T5 models

pytorch_bert_japanese

PytorchでBERTの日本語学習済みモデルを利用する

Laboro-BERT-Japanese

Laboro BERT Japanese: Japanese BERT Pre-Trained With Web-Corpus

RoBERTa-japanese

Japanese BERT Pretrained Model

aMLP-japanese

aMLP Transformer Model for Japanese

bert-japanese-aozora

Japanese BERT trained on Aozora Bunko and Wikipedia, pre-tokenized by MeCab with UniDic & SudachiPy

sbert-ja

Code to train Sentence BERT Japanese model for Hugging Face Model Hub

BERT-Japan-vaccination

Official fine-tuning code for "Emotion Analysis of Japanese Tweets and Comparison to Vaccinations in Japan"

gpt2-japanese

Japanese GPT2 Generation Model

text2text-japanese

gpt-2 based text2text conversion model

gpt-ja

GPT-2 Japanese model for HuggingFace's transformers

albert-japanese

BERT with SentencePiece for Japanese text.

ja_text_bert

日本語WikipediaコーパスでBERTのPre-Trainedモデルを生成するためのリポジトリ

DistilBERT-base-jp

A Japanese DistilBERT pretrained model, which was trained on Wikipedia.

bert

This repository provides snippets to use RoBERTa pre-trained on Japanese corpus. Our dataset consists of Japanese Wikipedia and web-scrolled articles, 25GB in total. The released model is built based on that from HuggingFace.

Laboro-DistilBERT-Japanese

Laboro DistilBERT Japanese

luke

LUKE -- Language Understanding with Knowledge-based Embeddings

GPTSAN

General-purpose Swich transformer based Japanese language mode

In 2 lists

AcademicBART

We pretrained a BART-based Japanese masked language model on paper abstracts from the academic database CiNii Articles

AcademicRoBERTa

We pretrained a RoBERTa-based Japanese masked language model on paper abstracts from the academic database CiNii Articles.

LINE-DistilBERT-Japanese

DistilBERT model pre-trained on 131 GB of Japanese web text. The teacher model is BERT-base that built in-house at LINE.

Japanese-Alpaca-LoRA

日本語に翻訳したStanford Alpacaのデータセットを用いてLLaMAをファインチューニングし作成したLow-Rank AdapterのリンクとGenerateサンプルコード

albert-japanese-tinysegmenter

Pretrained models, codes and guidances to pretrain official ALBERT(https://github.com/google-research/albert) on Japanese Wikipedia Resources

japanese-llama-experiment

Japanese LLaMa experiment

easylightchatassistant

EasyLightChatAssistant は軽量で検閲や規制のないローカル日本語モデルのLightChatAssistant を、KoboldCpp で簡単にお試しする環境です。

ChatGPT

VRChatGPT

ChatGPTを使ってVRChat上でお喋り出来るようにするプログラム。

AITuberDegikkoMirii

AITuberの基礎となる部分を開発しています

wanna

Shell command launcher with natural language

In 2 lists

ChatdollKit

ChatdollKit enables you to make your 3D model into a chatbot

In 5 listsDetails

ChuanhuChatGPTJapanese

GUI for ChatGPT API For Japanese

AISisterAIChan

ChatGPT3.5を搭載した伺かゴースト「AI妹アイちゃん」です。利用には別途ChatGPTのAPIキーが必要です。

In 2 lists

vrchatbot

VRChatにAI Botを作るためのリポジトリ

In 2 lists

gptuber-by-langchain

GPTがYouTuberをやります

In 2 lists

openai-chatfriend

A chatbox application built using Nuxt 3 powered by Open AI Text completion endpoint. You can select different personality of your AI friend. The default will respond in Japanese. You can use this app to practice your Nihongo skills!

chrome-ext-translate-to-hiragana-with-chatgpt

This Chrome extension can translate selected Japanese text to Hiragana by using ChatGPT.

azure-search-openai-demo

このサンプルでは、Retrieval Augmented Generation パターンを使用して、独自のデータに対してChatGPT のような体験を作成するためのいくつかのアプローチを示しています。

chatvrm

ChatVRMはブラウザで簡単に3Dキャラクターと会話ができるデモアプリケーションです。

sftly-replace

A Chrome extention to replace the selected text softly

In 2 lists

summarize_arxv

Summarize arXiv paper with figures

aiavatarkit

Building AI-based conversational avatars lightning fast

In 2 lists

jp-azureopenai-samples

Azure OpenAIを活用したアプリケーション実装のリファレンスを目的として、アプリのサンプル(リファレンスアーキテクチャ、サンプルコードとデプロイ手順)を無償提供しています。

In 2 lists

character_chat

OpenAIのAPIを利用して、設定したキャラクターと日本語で会話するチャットスクリプトです。

chatgpt-slackbot

OpenAIのChatGPT APIをSlack上で利用するためのSlackbotスクリプト (日本語での利用が前提)

In 2 lists

chatgpt-prompt-sample-japanese

ChatGPT の Prompt のサンプルです。

In 2 lists

kanji-flashcard-app-gpt4

A Japanese Kanji Flashcard App built using Python and Langchain, enhanced with the intelligence of GPT-4.

IgakuQA

Evaluating GPT-4 and ChatGPT on Japanese Medical Licensing Examinations

japagen

日本語タスクにおけるLLMを用いた疑似学習データ生成の検討

generativeai-prompt-sample-japanese

ChatGPTやCopilotなど各種生成AI用の「日本語]の Prompt のサンプル

In 2 lists

Dictionary and IME

mecab-ipadic-neologd

Neologism dictionary based on the language resources on the Web for mecab-ipadic

tdmelodic

A Japanese accent dictionary generator

jamdict

Python 3 library for manipulating Jim Breen's JMdict, KanjiDic2, JMnedict and kanji-radical mappings

unidic-py

Unidic packaged for installation via pip.

Japanese-Company-Lexicon

Japanese Company Lexicon (JCLdic)

manbyo-sudachi

Sudachi向け万病辞書

jawiki-kana-kanji-dict

Generate SKK/MeCab dictionary from Wikipedia(Japanese edition)

JIWC-Dictionary

dictionary to find emotion related to text

JumanDIC

This repository contains source dictionary files to build dictionaries for JUMAN and Juman++.

ipadic-py

IPAdic packaged for easy use from Python.

unidic-lite

A small version of UniDic for easy pip installs.

emoji-ime-dictionary

日本語で絵文字入力をするための IME 追加辞書 orange_book Google 日本語入力などで日本語から絵文字への変換を可能にする IME 拡張辞書

google-ime-dictionary

日英変換・英語略語展開のための IME 追加辞書 orange_book 日本語から英語への和英変換や英語略語の展開を Google 日本語入力や ATOK などで可能にする IME 拡張辞書

dic-nico-intersection-pixiv

ニコニコ大百科とピクシブ百科事典の共通部分のIME辞書

google-ime-user-dictionary-ja-en

GoogleIME用カタカナ語辞書プロジェクトのアーカイブです。Project archive of Google IME user dictionary from Katakana word ( Japanese loanword ) to English.

emoticon

Google日本語入力の顔文字辞書∩(,,Ò‿Ó,,)∩

mecab-mozcdic

open source mozc dictionaryをMeCab辞書のフォーマットに変換したものです。

denonbu-ime-dic

電音IME: Microsoft IMEなどで利用することを想定した「電音部」関連用語の辞書

nijisanji-ime-dic

Microsoft IMEなどで利用することを想定した「にじさんじ」関連用語の用語辞書です。

pokemon-ime-dic

Microsoft IMEなどで利用することを想定した、現状判明している全てのポケモンの名前を網羅した用語辞書です。

EJDict

English-Japanese Dictionary data (Public Domain) EJDict-hand

Ayashiy-Nipongo-Dic

贵樣ばこゐ辞畫を使て正レい日本语を使ラことが出來ゑ。

genshin-dict

Windows/macOSで使える原神の単語辞書です

jmdict-simplified

JMdict and JMnedict in JSON format

mozcdict-ext

Convert external words into Mozc system dictionary

mh-dict-jp

MonsterHunterのユーザー辞書を作りたい…

mecab-unidic-neologd

Neologism dictionary based on the language resources on the Web for mecab-unidic

hololive-dictionary

ホロライブ(ホロライブプロダクション)に関する辞書ファイルです。./dictionary フォルダ内のテキストファイルを使って、IMEに単語を追加できます。詳細はREADME.mdをご覧ください。

jmdict-yomitan

JMdict, JMnedict, KANJIDIC for Yomitan/Yomichan.

yomichan-jlpt-vocab

JLPT level tags for words in Yomichan

Jitendex

A free and openly licensed Japanese-to-English dictionary compatible with multiple dictionary clients

jiten

japanese android/cli/web dictionary based on jmdict/kanjidic — 日本語 辞典 和英辞典 漢英字典 和独辞典 和蘭辞典

pixiv-yomitan

Pixiv Encyclopedia Dictionary for Yomitan

uchinaaguchi_dict

うちなーぐち辞典(沖縄語辞典)

yomitan-dictionaries

Japanese and Chinese dictionaries for Yomitan.

mouse_over_dictionary

マウスオーバーした単語を自動で読み取る汎用辞書ツール

jisyo

かな漢字変換エンジン SKKのための新しい辞書形式

skk-jisyo.emoji-ja

日本語の読みから Emoji に変換するための SKK 辞書 😂

anthy

Anthy is a kana-kanji conversion engine for Japanese. It converts roma-ji to kana, and the kana text to a mixed kana and kanji.

aws_dic_for_google_ime

AWSサービス名のGoogle日本語入力向けの辞書

cl-skkserv

Common LispによるSKK辞書サーバーとその拡張

anthy

Anthy maintenance

anthy-unicode

Anthy Unicode - Another Anthy

azooKey

azooKey is an open-source Japanese keyboard for iPhone and iPad, written in Swift and powered by its own kana-kanji conversion engine. It provides live conversion, flexible key layouts, and a clean SwiftUI interface for a smooth typing experience.

azooKey-Desktop

azooKey-Desktop is an open-source Japanese input method for macOS, written in Swift and powered by the Zenzai neural kana-kanji converter. It provides live conversion, optional LLM-based “Magic Conversions”, and Tuner-backed personalization for a smooth, desktop typing experience.

fcitx5-hazkey

Japanese input method for fcitx5, powered by azooKey engine

mozcdic-ut-place-names

Mozc UT Place Name Dictionary is a dictionary converted from the Japan Post's ZIP code data for Mozc.

AzooKeyKanaKanjiConverter

Kana-Kanji Conversion Module written in Swift, supporting Neural Kana-Kanji Conversion and other cool features.

libkkc

Japanese Kana Kanji conversion input method library

libskk

Japanese SKK input method library

cjkvi-dict

漢字データベースの辞書関連データ

wlsp-classical

古典日本語の分類語彙表データ

kanji-dict

漢字の書き順(筆順)・読み方・画数・部首・用例・成り立ちを調べるための漢字辞書です。Unicode 15.1 のすべての漢字 98,682字を収録しています。

Kaomoji_proj

(๑ ᴖ ᴑ ᴖ ๑)みょんかおもじ(旧Kaomoji_proj)はMicrosoft社の入力ソフト、Microsoft IME向けの顔文字の辞書を作成するプロジェクトです。

kotlin-kana-kanji-converter

Kotlin かな漢字変換プログラム

alfred-japanese-dictionary

Japanese-English Dictionary using jisho.org with audio, csv export of entries, and preview of dictionary sites.

ichiran

Linguistic tools for texts in Japanese language

mikan

A Japanese input method.

colloquial-kansai-dictionary

A quick reference for the material taught in Colloquial Kansai Japanese.

jisho-open

Web frontend for the JMdict Japanese-English dictionary project, with study list support!

macskk

Yet Another macOS SKK Input Method

nandoku

難読漢字を学年別にまとめた辞書です。

japanese_android_ime

A FOSS Japanese IME for Android

anthywl

Japanese input method for Sway using libanthy

sekka

Yet another Japanese Input Method inspired by SKK.

sumibi

Japanese input method powered by ChatGPT API

jinmei-dict

辞書データから人名だけを抜き出し、読み仮名(カタカナ)をキーとして、候補となる書き文字をリストで保持するようなJSON形式に整形しています。

japanesearabic

JapaneseArabic Dictionary (日本語・アラビア語辞書) قاموس اللغة اليابانية والعربية (Yomitan)

o-dic

沖縄辞書

skk-emoji-jisyo

SKK 絵文字辞書

mozcdic-ut-personal-names

A personal name dictionary for Mozc.

mozcdic-ut-sudachidict

A dictionary converted from SudachiDict for Mozc.

nihongo

japanese language data and dictionary

kagome-dict

Dictionary Library for Kagome v2

canna

Canna Japanese input system

kansai-accent-dictionary

京阪式アクセント(関西弁)辞書 - 4,615語を収録した日本語方言アクセント辞書

jitendex

A free, offline, and openly licensed Japanese-to-English dictionary. Updates monthly!

karukan

Japanese Input Method System for Linux, Neural Kana-Kanji Conversion Engine + fcitx5 IME

shitto-mania-dic

嫉妬辞書(Shitto-Mania / Jealousy Dictionary)

dvorakjp-roman-table

azooKey, Google 日本語入力用 DvorakJP ローマ字テーブル / DvorakJP Roman Table for azooKey, Google Japanese Input

jmdict-fst

Fast JMdict lookup engine with FST-based exact/prefix/fuzzy/gloss search, deinflection, Rust core, and Swift/Kotlin/Flutter bindings.

mzimeja

MZ-IME Japanese Input for Windows

japanesekeyboard

スミレ - 完全オフラインの日本語キーボードアプリ

rakukan

ローカルLLMを利用した、Windows 向け日本語 IMEgit

JMdictSQLite

SQLite database for JMdict and Kanjidic, a Japanese-English dictionary. Automatic daily updates.

Corpus >Part-of-speech tagging / Named entity recognition

ner-wikipedia-dataset

Wikipediaを用いた日本語の固有表現抽出データセット

IOB2Corpus

Japanese IOB2 tagged corpus for Named Entity Recognition.

TwitterCorpus

首都大日本語 Twitter コーパス

UD_Japanese-PUD

Parallel Universal Dependencies.

UD_Japanese-GSD

Japanese data from the Google UDT 2.0.

KWDLC

Kyoto University Web Document Leads Corpus

AnnotatedFKCCorpus

Annotated Fuman Kaitori Center Corpus

UD_Japanese-GSDLUW

Long-unit-word version of UD_Japanese-GSD

ud_japanese-bccwj

This Universal Dependencies (UD) Japanese treebank is based on the definition of UD Japanese convention described in the UD documentation.

anthy

Anthy is a kana-kanji conversion engine for Japanese. It converts roma-ji to kana, and the kana text to a mixed kana and kanji.

Corpus >Parallel corpus

small_parallel_enja

50k English-Japanese Parallel Corpus for Machine Translation Benchmark.

Web-Crawled-Corpus-for-Japanese-Chinese-NMT

A Web Crawled Corpus for Japanese-Chinese NMT

CourseraParallelCorpusMining

Coursera Corpus Mining and Multistage Fine-Tuning for Improving Lectures Translation

JESC

A large parallel corpus of English and Japanese

AMI-Meeting-Parallel-Corpus

AMI Meeting Parallel Corpus

giant_ja-en_parallel_corpus

This directory includes a giant Japanese-English subtitle corpus. The raw data comes from the Stanford’s JESC project.

jesc_small

Small Japanese-English Subtitle Corpus

graded-enja-corpus

禁止用語や単語レベルを考慮した日英対訳コーパスです。

cjk-compsci-terms

CJK computer science terms comparison / 中日韓電腦科學術語對照 / 日中韓のコンピュータ科学の用語対照 / 한·중·일 전산학 용어 대조

Laboro-ParaCorpus

Scripts for creating a Japanese-English parallel corpus and training NMT models

google-vs-deepl-je

google-vs-deepl-je

matcha

訪日観光客向けメディアMATCHAの記事から、日本語のテキスト平易化のためのデータセットを構築しました。

en-ja-el

EnJaEL: En-Ja Parallel Entity Linking Dataset (Version 1.0)

Corpus >Dialog corpus

JMRD

Japanese Movie Recommendation Dialogue dataset

open2ch-dialogue-corpus

おーぷん2ちゃんねるをクロールして作成した対話コーパス

BSD

The Business Scene Dialogue corpus

asdc

Accommodation Search Dialog Corpus (宿泊施設探索対話コーパス)

japanese-corpus

日本語の対話データ for seq2seq etc

BPersona-chat

This repository contains the Japanese–English bilingual chat corpus BPersona-chat published in the paper Chat Translation Error Detection for Assisting Cross-lingual Communications at AACL-IJCNLP 2022's Workshop Eval4NLP 2022.

japanese-daily-dialogue

Japanese Daily Dialogue, or 日本語日常対話コーパス in Japanese, is a high-quality multi-turn dialogue dataset containing daily conversations on five topics: dailylife, school, travel, health, and entertainment.

llm-japanese-dataset

LLM構築用の日本語チャットデータセット

kokorochat

ロールプレイで収集した日本語のカウンセリング対話データセット

JMultiWOZ-TC

マルチターン対話でのエージェントのfunction calling評価

HOTATE

本音・建前付き日本語対話データセット

ETCDataset

Emotion Transcription in Conversation Dataset は,対話中の各発話に対して話者自身が記述した心情文を含む,約1,000 件の対話からなる日本語対話データセットです.

Show 189 items

jrte-corpus

Japanese Realistic Textual Entailment Corpus (NLP 2020, LREC 2020)

kanji-data

A JSON kanji dataset with updated JLPT levels and WaniKani information

JapaneseWordSimilarityDataset

Japanese Word Similarity Dataset

simple-jppdb

A paraphrase database for Japanese text simplification

chABSA-dataset

chakki's Aspect-Based Sentiment Analysis dataset

JaQuAD

JaQuAD: Japanese Question Answering Dataset for Machine Reading Comprehension (2022, Skelter Labs)

JaNLI

Japanese Adversarial Natural Language Inference Dataset

ebe-dataset

Evidence-based Explanation Dataset (AACL-IJCNLP 2020)

emoji-ja

UNICODE絵文字の日本語読み/キーワード/分類辞書

nayose-wikipedia-ja

Wikipediaから作成した日本語名寄せデータセット

ja.text8

Japanese text8 corpus for word embedding.

ThreeLineSummaryDataset

3行要約データセット

japanese

This repo contains a list of the 44,998 most common Japanese words in order of frequency, as determined by the University of Leeds Corpus.

kanji-frequency

Kanji usage frequency data collected from various sources

TEDxJP-10K

TEDxJP-10K ASR Evaluation Dataset

CoARiJ

Corpus of Annual Reports in Japan

technological-book-corpus-ja

日本語で書かれた技術書を収集した生コーパス/ツール

ita-corpus-chuwa

Chunked word annotation for ITA corpus

wikipedia-utils

Utility scripts for preprocessing Wikipedia texts for NLP

inappropriate-words-ja

日本語における不適切表現を収集します。自然言語処理の時のデータクリーニング用等に使えると思います。

house-of-councillors

参議院の公式ウェブサイトから会派、議員、議案、質問主意書のデータを整理しました。

house-of-representatives

国会議案データベース:衆議院

STAIR-captions

STAIR captions: large-scale Japanese image caption dataset

Winograd-Schema-Challenge-Ja

Japanese Translation of Winograd Schema Challenge

speechBSD

An extension of the BSD corpus with audio and speaker attribute information

ita-corpus

ITAコーパスの文章リスト

rohan4600

モーラバランス型日本語コーパス

anlp-jp-history

言語処理学会年次大会講演の全リスト・機械可読版など

keigo_transfer_task

敬語変換タスクにおける評価用データセット

loanwords_gairaigo

English loanwords in Japanese

jawikicorpus

Japanese-Wikipedia Wikification Corpus

GeneralPolicySpeechOfPrimeMinisterOfJapan

This is the corpus of Japanese Text that general policy speech of prime minister of Japan

wrime

WRIME: 主観と客観の感情分析データセット

jtubespeech

JTubeSpeech: Corpus of Japanese speech collected from YouTube

WikipediaWordFrequencyList

日本語Wikipediaで使用される頻出単語のリスト

kokkosho_data

車両不具合情報に関するデータセット

pdmocrdataset-part1

デジタル化資料OCRテキスト化事業において作成されたOCR学習用データセット

huriganacorpus-ndlbib

全国書誌データから作成した振り仮名のデータセット

jvs_hiho

JVS (Japanese versatile speech) コーパスの自作のラベル

hirakanadic

Allows Sudachi to normalize from hiragana to katakana from any compound word list

animedb

約100年に渡るアニメ作品リストデータベース

In 2 lists

security_words

サイバーセキュリティに関連する公的な組織の日英対応

Data-on-Japanese-Diet-Members

日本の国会議員のデータ

honkoku-data

歴史資料の市民参加型翻刻プラットフォーム「みんなで翻刻」のテキストデータ置き場です。 / Transcription texts created on Minna de Honkoku (https://honkoku.org), a crowdsourced transcription platform for historical Japanese documents.

engineer-vocabulary-list

Engineer Vocabulary List in Japanese/English

JSICK

Japanese Sentences Involving Compositional Knowledge (JSICK) Dataset/JSICK-stress Test Set

phishurl-list

Phishing URL dataset from JPCERT/CC

jcms

A Japanese Corpus of Many Specialized Domains (JCMS)

aozorabunko_text

text-only archives of www.aozora.gr.jp

topokanji

Topologically ordered lists of kanji for effective learning

isbn4groups

ISBN-13における日本語での出版物 (978-4-XXXXXXXXX) に関するデータ等

NMeCab

NMeCab: About Japanese morphological analyzer on .NET

ndlngramdata

デジタル化資料から作成したOCRテキストデータのngram頻度統計情報のデータセット

ndlngramviewer_v2

2023年1月にリニューアルしたNDL Ngram Viewerのソースコード等一式

data_set

法律・判例関係のデータセット

huggingface-datasets_wrime

WRIME for huggingface datasets

ndl-minhon-ocrdataset

NDL古典籍OCR学習用データセット(みんなで翻刻加工データ)

PAX_SAPIENTICA

GIS & Archaeological Simulator. 2023 in development.

j-liwc2015

Japanese version of LIWC2015

huggingface-datasets_livedoor-news-corpus

Japanese Livedoor news corpus for huggingface datasets

huggingface-datasets_JGLUE

JGLUE: Japanese General Language Understanding Evaluation for huggingface datasets

commonsense-moral-ja

JCommonsenseMorality is a dataset created through crowdsourcing that reflects the commonsense morality of Japanese annotators.

comet-atomic-ja

COMET-ATOMIC ja

dcsg-ja

Dialogue Commonsense Graph in Japanese

japanese-toxic-dataset

"Proposal and Evaluation of Japanese Toxicity Schema" provides a schema and dataset for toxicity in the Japanese language.

camera

CAMERA (CyberAgent Multimodal Evaluation for Ad Text GeneRAtion) is the Japanese ad text generation dataset.

Japanese-Fakenews-Dataset

日本語フェイクニュースデータセット

copa-japanese

COPA Dataset in Japanese

WLSP-familiarity

Word Familiarity Rate for 'Word List by Semantic Principles (WLSP)'

ProSub

A cross-linguistic study of pronoun substitutes and address terms

ramendb

なんとかデータベース( https://supleks.jp/ )からのスクレイピングツールと収集データ

huggingface-datasets_CAMERA

CAMERA (CyberAgent Multimodal Evaluation for Ad Text GeneRAtion) for huggingface datasets

FactCheckSentenceNLI-FCSNLI-

FactCheckSentenceNLIデータセット

databricks-dolly-15k-ja

databricks/dolly-v2-12b の学習データに使用されたdatabricks-dolly-15k.jsonl を日本語に翻訳したデータセットになります。

EaST-MELD

EaST-MELD is an English-Japanese dataset for emotion-aware speech translation based on MELD.

meconaudio

Mecon Audio(Medical Conference Audio)は厚生労働省主催の先進医療会議の議事録の読み上げデータセットです。

japanese-addresses

全国の町丁目レベル(277,191件)の住所データのオープンデータ

aozorasearch

The full-text search system for Aozora Bunko by Groonga. 青空文庫全文検索ライブラリ兼Webアプリ。

llm-jp-corpus

This repository contains scripts to reproduce the LLM-jp corpus.

alpaca_ja

alpacaデータセットを日本語化したものです

instruction_ja

Japanese instruction data (日本語指示データ)

japanese-family-names

Top 5000 Japanese family names, with readings, ordered by frequency.

kanji-data-media

Japanese language data on kanji, radicals, media files, fonts and related resources from Kanji alive

reazonspeech

Construct large-scale Japanese audio corpus at home

huriganacorpus-aozora

青空文庫及びサピエの点字データから作成した振り仮名のデータセット

koniwa

An open collection of annotated voices in Japanese language

JMMLU

日本語マルチタスク言語理解ベンチマーク Japanese Massive Multitask Language Understanding Benchmark

hurigana-speech-corpus-aozora

青空文庫振り仮名注釈付き音声コーパスのデータセット

jqara

JQaRA: Japanese Question Answering with Retrieval Augmentation - 検索拡張(RAG)評価のための日本語Q&Aデータセット

jemhopqa

JEMHopQA (Japanese Explainable Multi-hop Question Answering) is a Japanese multi-hop QA dataset that can evaluate internal reasoning.

jacred

Repository for Japanese Document-level Relation Extraction Dataset (plan to be released in March).

jades

JADES is a dataset for text simplification in Japanese, described in "JADES: New Text Simplification Dataset in Japanese Targeted at Non-Native Speakers" (the paper will be available soon).

do-not-answer-ja

2023年8月にメルボルン大学から公開された安全性評価データセット『Do-Not-Answer』を日本語LLMの評価においても使用できるように日本語に自動翻訳し、さらに日本文化も考慮して修正したデータセット。

oasst1-89k-ja

OpenAssistant のオープンソースデータ OASST1 を日本語に翻訳したデータセットになります。

jacwir

JaCWIR: Japanese Casual Web IR - 日本語情報検索評価のための小規模でカジュアルなWebタイトルと概要のデータセット

japanese-technical-dict

日本語学習者のための科学技術業界でよく使われる片仮名と元の単語対照表

j-unimorph

Dataset of UniMorph in Japanese

GazeVQA

Dataset for the LREC-COLING 2024 paper "A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions"

J-CRe3

Code for J-CRe3 experiments (Ueda et al., LREC-COLING, 2024)

jmed-llm

JMED-LLM: Japanese Medical Evaluation Dataset for Large Language Models

lawtext

Plain text format for Japanese law

pdmocrdataset-part2

OCR処理プログラム研究開発事業において作成されたOCR学習用データセット

japanesetopicwsd

話題に基づく語義曖昧性解消評価セット

temporalNLI_dataset

Jamp: Controlled Japanese Temporal Inference Dataset for Evaluating Generalization Capacity of Language Models

JSeM

Japanese semantic test suite (FraCaS counterpart and extensions)

niilc-qa

NIILC QA data

chain-of-thought-ja-dataset

Dataset of paper "Verification of Chain-of-Thought Prompting in Japanese"

WikipediaAnnotatedCorpus

This is a Japanese text corpus that consists of Wikipedia articles with various linguistic annotations.

elaws-history

e-Gov 法令検索で配布されている「全ての法令データ」を定期的にダウンロードし、アーカイブしています

Japanese-RP-Bench

Japanese-RP-BenchはLLMの日本語ロールプレイ能力を測定するためのベンチマークです。

hdic

HDIC : Integrated Database of Hanzi Dictionaries in Early Japan

awesome-japan-opendata

Awesome Japan Open Data - 日本のオープンデータ情報一覧・まとめ

kanji-data

常用漢字表他、漢字に関するデータ

openchj-genji

「源氏物語」形態論情報データ

AdParaphrase

This repository contains data for our paper "AdParaphrase: Paraphrase Dataset for Analyzing Linguistic Features toward Generating Attractive Ad Texts".

Jamp_sp

アスペクトを考慮した日本語時間推論データセットの構築(Jamp_sp: Controlled Japanese Temporal Inference Dataset Considering Aspect)

jnli-neg

否定理解能力を評価するための日本語言語推論データセット JNLI-Neg の公開用リポジトリです。

swallow-corpus

This repository provides Python implementation for building Swallow Corpus Version 1, a large Japanese web corpus (Okazaki et al., 2024), from Common Crawl archives.

jalecon

A Dataset of Japanese Lexical Complexity for Non-Native Readers

multils-japanese

MultiLS-Japanese Lexical Complexity Prediction and Lexical Simplification Dataset for Japanese: annotator profiles, unaggregated annotation, and annotatation guidelines.

nwjc

NINJAL Web Japanese Corpus

open-mantra-dataset

Dataset introduced in the paper "Towards Fully Automated Manga Translation" presented in AAAI21

public-annotations

Various annotations of Manga109 dataset

gimei

random Japanese name and address generator

safety-boundary-test

日本語言語モデルの安全性の振る舞いを評価するテストセット

j-ono-data

A simple, open-source collection of Japanese onomatopoeic and mimetic sound words in JSON format. With manga samples.

kanji

List of japanese kanji radicals to learn

jethics

日本語道徳理解度評価用データセットJETHICSの概説ページ (to be update)

waon

WAON: Large-Scale and High-Quality Japanese Image-Text Dataset for Vision-Language Models

kuci

Kyoto University Commonsense Inference dataset (KUCI)

japanese-address-testdata

解析が難しい日本の住所のテストデータセット

jlpt-word-list

Japanese word list from JLPT vocabulary

hiragana_mojigazo

文字画像データセット(平仮名73文字版)

lawqa_jp

日本の法令に関する多肢選択式QAデータセット

yjcaptions

YJ Captions 26k Dataset

ja-vg-vqa

Japanese Visual Genome VQA dataset

lawhub

Repository to track Japanese Law in text format

jconj

A table-based Japanese word conjugator

extract_jawp_names

Extracts personal names in Wikipedia Japanese.

cejc_yomichan_freq_dict

Frequency dictionary for yomichan based on the Corpus of Everyday Japanese Conversation dataset

wikidict-ja

Wikipedia Bilingual Reference Data (Japanese)

ajimee-bench

AJIMEE-Bench (Advanced Japanese IME Evaluation Benchmark)

j-spaw

J-SpAW: Japanese speech corpus for speaker verification and anti-spoofing

camera3

CAMERA3: An Evaluation Dataset for Controllable Ad Text Generation in Japanese

jgpqa

Japanese translation of the GPQA dataset

tanaka-corpus-plus

Tanaka Corpus のノイズを除去しています。

emotioncorpusjapanesetokushimaa2lab

Japanese emotion corpus Tokushima Univ. A-2 Lab.

osworld-jp

言語を考慮した評価のための、日本語版コンピュータユースベンチマーク

quasi_japanese_reviews

Quasi Japanese Reviews (擬似レビューデータ)

psychiatry-clinical-notes

精神科初診カルテ作成アンケート データセット

merged-town-names

市町村合併などにより消滅した旧地名と新地名の対応表

japanesetextemoticondata

Japanese text-emoticon data.

mishearing-corpus

聞き間違えコーパス︱CSV+Table Schema で約 1 万件を管理し、VS Code+pre-commit+Frictionless+GitHub Actions で自動検証を行う日本語データセット

kotowaza

Structured JSON dataset of Japanese proverbs (kotowaza) with meanings in Indonesian & English, examples, JLPT levels, and tags.

selective-rag-kasensabo

建設の技術基準に関する質問の専門性粒度(細かい/粗い)を96%正確に自動判定し、最適なRAGシステム(ColBERT/Naive)を選択する実用的なAgentic RAGシステムのMVPです。2025年11月に公開された河川砂防ダムの技術基準を対象に4つのRAGシステムを構築し、専門性の粒度が異なる200問の質問に対して、精度と速度を比較した。

jmle2026-bench

LLM benchmark on the 120th Japanese Medical Licensing Examination (Feb 7-8, 2026)

JSTS-Neg

否定理解能力を評価するための日本語意味的類似度計算データセット JSTS-Neg の公開用リポジトリです。 JSTS-Neg は、JGLUE に含まれる言語推論データセット JSTS を拡張して作成しました。

business-slide-questions

このリポジトリでは、ビジネス資料(スライド)を対象とした Visual Question Answering (VQA) ベンチマーク「BusinessSlideVQA」を提供しています。

WLSP-antonym

Antonym relations for 'Word List by Semantic Principles (WLSP)'

YouCook2-JP

Japanese translation of the YouCook2 dataset.

E2U

つたわる化に関するデータ

annotation-2025

このリポジトリは,テキストの「解釈」を人手とLLM出力で比較できるデータを公開するためのものです.

jhpt

歴史的日本語資料の原文テキストと,現代語訳(参照訳)テキストをセグメント単位で対応付けた対訳データセットです.詳細は論文を参照ください.

JBE-QA

Japanese Bar Exam QA

JMedWiC

マスク言語モデルを用いて擬似的な同義・非同義ペアを自動抽出し,人手による同義性アノテーションを通じてラベルを決定することで,日本語の医療分野における語義同一性判定データセットを構築しました.

Doppelganger-JC

This is a dataset benchmarking the misuse of cross-lingual homographs between Chinese and Japanese in LLMs.

modelvista-3lang

ソフトウェア図理解のためのVLM評価ベンチマーク(日本語・英語・韓国語対応)

japanese-hr-niah

日本語人事労務ドメインにおけるロングコンテキストLLMの性能評価ベンチマーク

nijl-manyoshutei

本リポジトリでは、関西大学所蔵廣瀬本万葉集のTEI/XMLデータ等をCC-BYライセンスのもとで公開しています。

kamuskita

マレー語勉強会で作っているオープンなマレー語・日本語辞典『みんなのマレー語辞典』

japanese-llm-benchmark

A benchmark tool for evaluating Japanese language capabilities of various LLMs.

EDINET-Bench

ICLR 2026 Evaluating the performance of LLMs on Japanese challenging financial tasks.

LCTG-Bench

LCTG Bench: LLM Controlled Text Generation Benchmark

Kokoro-Speech-Dataset

A public domain single speaker Japanese speech dataset

LookVQA

A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions (LREC-COLING 2024)

JTruthfulQA

JTruthfulQA is a Japanese version of TruthfulQA (Lin+, 2022). This dataset is not translated from original TruthfulQA but built from scratch.

japanese-dataset-for-automated-fact-checking

Japanese Dataset for Automated Fact-Checking: JAD-AFC

llm-jp-longbench

日本語版longbench作成のため、jemhopデータセットを活用。

doctrine-corpus

Bilingual (Japanese + English) judgment-eliciting Q&A corpus (851 examples) encoding documented research-program decisions, with per-example metadata (CC0, Hugging Face mirror available)

medLLM_QA_benchmark

Trilingual (English, Japanese, Chinese) QA benchmark for medical LLM

kaomoji-data

約1,700種類の日本語顔文字を全件、かな読み(擬態語・キーワード)・タグ・カテゴリでアノテーションした構造化JSONデータセット。多カテゴリ対応・カテゴリ別ファイル分割、MITライセンス。

jvs_nonpara_kana

Katakana annotation of JVS nonpara corpus for G2P evaluation

jlpt-kanji-dictionary

Structured Japanese Kanji and Vocabulary JSON datasets organized by JLPT level with English and Russian translations — ready for use in language learning apps, NLP, and kanji study tools.

pfgen-bench

Preferred Generation Benchmark

j-tau-bench

J-tau: A Japanese tau-bench for Benchmarking Tool-Agent-User Interaction in Real-World Domains

bbh-ja

Japanese Translation of BIG-Bench-Hard (https://github.com/suzgunmirac/BIG-Bench-Hard/)

jfbench

JFBench: Japanese instruction Following Benchmark

aica-corpus

AIキャラクター・フィラー・笑い声・感情表現に特化した日本語TTSコーパス(CC0)

adlib

ADLIB: Japanese ASR benchmark framework with language-aware evaluationb

Tutorial

spacy_tutorial

spaCy tutorial in English and Japanese. spacy-transformers, BERT, GiNZA.

fastTextJapaneseTutorial

Tutorial to train fastText with Japanese corpus

allennlp-NER-ja

AllenNLP-NER-ja: AllenNLP による日本語を対象とした固有表現抽出

chariot-PyTorch-Japanese-text-classification

Experiment for Japanese Text classification using chariot and PyTorch

ginza-examples

日本語NLPライブラリGiNZAのすゝめ

DocumentClassificationUsingBERT-Japanese

DocumentClassificationUsingBERT-Japanese

BERT_Japanese_Google_Colaboratory

Google Colaboratoryで日本語のBERTを動かす方法です。

bert-book

「BERTによる自然言語処理入門: Transformersを使った実践プログラミング」サポートページ

janome-tutorial

Janome を使ったテキストマイニング入門チュートリアルです。

handson-language-models

日本語の言語モデルのハンズオン資料です

JapaneseNLI

Google Colabで日本語テキスト推論を試す

deep-learning-with-pytorch-ja

deep-learning-with-pytorchの日本語版repositoryです。

bert-classification-tutorial

【2023年版】BERTによるテキスト分類

python-nlp-book

ディープラーニングによる自然言語処理(共立出版)のサポートページです

llm-book

「大規模言語モデル入門」(技術評論社, 2023)のGitHubリポジトリ

nlp2024-tutorial-3

NLP2024 チュートリアル3 作って学ぶ日本語大規模言語モデル - 環境構築手順とソースコード

japanese-ir-tutorial

日本語情報検索チュートリアル

nlpbook

「自然言語処理の教科書」サポートサイト

kantan-regex-book

作って学ぶ正規表現エンジン

bert-classification-tutorial-2024

【2024年版】BERTによるテキスト分類

Gemma2_2b_Japanese_finetuning_colab.ipynb

Fine-Tuning Google Gemma for Japanese Instructions

nlp100v2020

「言語処理100本ノック 2020」をPythonで解く

textmining-ja

Rによる自然言語処理・テキスト分析の練習

nlp2025-tutorial-2

NLP2025 のチュートリアル「地理情報と言語処理 実践入門」の資料とソースコード

nlp100v2025

「言語処理100本ノック 2025」をPythonで解く

topic-models-ao

『トピックモデル』(機械学習プロフェッショナルシリーズ)のノート

slp2025

音学シンポジウム2025チュートリアル「マルチモーダル大規模言語モデル入門」資料

book_impress_it-basic-education-ai

インプレス出版「IT基礎教養 自然言語処理&画像解析」

genai-agent-advanced-book

書籍「現場で活用するための生成AIエージェント実践入門」(講談社サイエンティフィック社)で利用されるソースコード

support-genai-book

原論文から解き明かす生成AI(技術評論社)のサポートページです

ir100

情報検索100本ノック

kaggle_llm_book

『Kaggle ではじめる大規模言語モデル入門 ~自然言語処理〈実践〉プログラミング~』のサポートサイト

nlp-lecture-keio

慶応義塾大学 理工学部 情報工学科 講義「自然言語処理」

llm-jp-4-cookbook

Example scripts for LLM-jp-4 models

ttslearn

ttslearn: Library for Pythonで学ぶ音声合成 (Text-to-speech with Python)

public-annotations

Various annotations of Manga109 dataset

Research summary

awesome-bert-japanese

A list of pre-trained BERT models for Japanese with word/subword tokenization + vocabulary construction algorithm information

GEC-Info-ja

文法誤り訂正に関する日本語文献を収集・分類するためのリポジトリ

dataset-list

lists of text corpus and more (mainly Japanese)

tuning_playbook_ja

ディープラーニングモデルの性能を体系的に最大化するためのプレイブック

japanese-pitch-accent-resources

Trying to consolidate japanese phonetic, and in particular pitch accent resources into one list

awesome-japanese-llm

オープンソースの日本語LLMまとめ

Reference

自然言語処理の餅屋

yasuokaの日記: 日本語係り受け解析器「2020年の総ざらえ」

yasuokaの日記: 日本語係り受け解析器「2021年の総ざらえ」

https://github.com/topics/japanese?l=python

https://github.com/topics/japanese-language?l=python

https://github.com/search?o=desc&q=corpus+japanese&s=&type=Repositories

https://paperswithcode.com/datasets?lang=japanese

awesome-bert-japanese

A list of pre-trained BERT models for Japanese with word/subword tokenization + vocabulary construction algorithm information

Awesome-Rust-MachineLearning-日本語向けのrustクレートや記事等をまとめたもの

大規模言語モデル入門Ⅱ 〜生成型LLMの実装と評価

See category
94

Awesome Rust

rust-unofficial/awesome-rust

A curated list of Rust code and resources.

Fresh★ 60k1833 entriesPushed today
94

Awesome Mac

jaywcjlove/awesome-mac

 This project is dedicated to collecting high-quality macOS software and organizing them systematically by different categories for easy search and use.

Fresh★ 115k1316 entriesPushed today
94

Awesome Python

vinta/awesome-python

The definitive list that answers "I want to do X in Python, which tool should I use?"

Fresh★ 324k504 entriesPushed today
94

Awesome C++

fffaraz/awesome-cpp

A curated list of awesome C++ (or C) frameworks, libraries, resources, and shiny things. Inspired by awesome-... stuff.

Fresh★ 74k1358 entriesPushed yesterday
94

Awesome Go

avelino/awesome-go

A curated list of awesome Go frameworks, libraries and software

Fresh★ 186k3119 entriesPushed yesterday
94

Awesome PHP

ziadoz/awesome-php

A curated list of amazingly awesome PHP libraries, resources and shiny things.

Fresh★ 33k87 entriesPushed 2 days ago