Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
-
Updated
Aug 15, 2026 - Python
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
Align Anything: Training All-modality Model with Feedback
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
[ICCV 2025] Official code of DeepMesh: Auto-Regressive Artist-mesh Creation with Reinforcement Learning
🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.
A curated list of papers on reinforcement learning for video generation
tensorflow를 사용하여 텍스트 전처리부터, Topic Models, BERT, GPT, LLM과 같은 최신 모델의 다운스트림 태스크들을 정리한 Deep Learning NLP 저장소입니다.
Train Large Language Models on MLX.
SiLLM simplifies the process of training and running Large Language Models (LLMs) on Apple Silicon by leveraging the MLX framework.
Zero-friction LLM fine-tuning skill for Claude Code, Gemini CLI & any ACP agent. Unsloth on NVIDIA · TRL+MPS/MLX on Apple Silicon. Automates env setup, LoRA training (SFT, DPO, GRPO, vision), post-hoc GRPO log diagnostics, evaluation, and export end-to-end. Part of the Gaslamp AI platform.
[CVPR 2025] Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization
[ICLR 2025] IterComp: Iterative Composition-Aware Feedback Learning from Model Gallery for Text-to-Image Generation
A framework for agentic tool use training with reinforcement learning
Add a description, image, and links to the dpo topic page so that developers can more easily learn about it.
To associate your repository with the dpo topic, visit your repo's landing page and select "manage topics."