Skip to content

Pull requests: huggingface/trl

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Reject unsupported train_dataset types in core trainers
#6493 opened Jul 21, 2026 by albertvillanova Member Loading…
Test train_dataset=None raises for core trainers
#6492 opened Jul 21, 2026 by albertvillanova Member Loading…
Record rollout traces to trackio
#6491 opened Jul 21, 2026 by AmineDiro Member Loading…
added step time metric to AsyncGRPOTrainer
#6490 opened Jul 21, 2026 by AmineDiro Member Loading…
Fix : queue wait time metric AsyncGRPOTrainer
#6489 opened Jul 21, 2026 by AmineDiro Member Loading…
docs: complete GRPOConfig loss_type help
#6477 opened Jul 20, 2026 by gowtham-sai-yadav Loading…
3 of 8 tasks
fix(dpo): keep apo_down on raw log-ratios for both terms
#6469 opened Jul 20, 2026 by Solaris-star Loading…
4 of 8 tasks
Raise a clear error for QLoRA + vllm_mode="server" instead of a cryptic vLLM crash
#6440 opened Jul 18, 2026 by DaoyuanLi2816 Contributor Loading…
3 of 4 tasks
Add privileged-context distillation to GOLD
#6437 opened Jul 17, 2026 by eshwanthkartitr Loading…
6 of 8 tasks
[SDPO] Port unscorable-mask reward defense from GRPO
#6429 opened Jul 17, 2026 by AUTHENSOR Loading…
5 of 8 tasks
Add LFM2 and LFM2.5 support and testing
#6428 opened Jul 16, 2026 by qgallouedec Member Loading…
Async grpo OpenEnv harness rollout
#6420 opened Jul 16, 2026 by AmineDiro Member Loading…
ProTip! Mix and match filters to narrow down what you’re looking for.