Skip to content

Entry

TRL

Appears in 8 awesome lists

Train transformer language models with reinforcement learning.

Open github.comhuggingface/trl

Found in these lists

When LLM Agents Meet Reinforcement Learning

Section: Base Framework · HuggingFace

FreshScore 81

Awesome LLMOps

Section: Foundation Model Fine Tuning · Train transformer language models with reinforcement learning.

ActiveScore 75

Awesome local LLM

Section: Training and Fine-tuning · train transformer language models with reinforcement learning

FreshScore 87

awesome-nlp

Section: Instruction Tuning and Preference Optimization · reference library for SFT, DPO, GRPO, and RLHF.

FreshScore 90

Awesome Open Source AI

Section: 7. Training & Fine-tuning Ecosystem · Official library for RLHF, SFT, DPO, ORPO.

FreshScore 89

Awesome Production Machine Learning

Section: Industry Strength Reinforcement Learning · Train transformer language models with reinforcement learning.

FreshScore 92

awesome-python

Section: Other · Train transformer language models with reinforcement learning.

FreshScore 81

Awesome Python Data Science

Section: Graph Machine Learning · Train transformer language models with reinforcement learning.

ActiveScore 71

PEFT

Parameter-Efficient Fine-Tuning (PEFT) methods enable efficient adaptation of pre-trained language models (PLMs) to various downstream applications without fine-tuning all the model's parameters.

In 7 listsDetails

slime

LLM post-training framework for RL Scaling from THUDM. Supports SFT and RL training with multi-turn compilation feedback, powering projects like TritonForge for automated GPU kernel generation. Apache 2.0 licensed.

In 5 listsDetails

Kiln

Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

In 5 listsDetails

OpenRLHF

an easy-to-use, high-performance open-source RLHF framework built on Ray, vLLM, ZeRO-3 and HuggingFace Transformers, designed to make RLHF training simple and accessible

In 4 listsDetails

LMFlow

Extensible toolkit for finetuning and inference of large foundation models. Features RAFT alignment algorithm and comprehensive model support. Apache 2.0 licensed.

In 4 listsDetails

alpaca-lora

Code for rproducing the Stanford Alpaca results using low-rank adaptation (LoRA).

In 4 listsDetails

RLinf

Scalable open-source RL infrastructure for post-training foundation models via reinforcement learning. Features M2Flow paradigm for embodied AI and agentic workflows with real-world robotics integrations. Apache 2.0 licensed.

In 3 lists

sentence-transformers

a Python library for using and training embedding and reranker models for applications like retrieval augmented generation, semantic search, and more

In 3 lists