-
Notifications
You must be signed in to change notification settings - Fork 2.9k
Pull requests: huggingface/trl
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Reject unsupported train_dataset types in core trainers
#6493
opened Jul 21, 2026 by
albertvillanova
Member
Loading…
Test train_dataset=None raises for core trainers
#6492
opened Jul 21, 2026 by
albertvillanova
Member
Loading…
Fix : queue wait time metric AsyncGRPOTrainer
#6489
opened Jul 21, 2026 by
AmineDiro
Member
Loading…
Add regression tests for evaluating init-time eval datasets after training with precomputed reference log-probs
#6488
opened Jul 21, 2026 by
albertvillanova
Member
Loading…
[DistillationTrainer refactor] Emit prompt_ids / prompt_mask / completion_ids
#6487
opened Jul 21, 2026 by
qgallouedec
Member
Loading…
fix(sft): support functools.partial model.forward in chunked CE patch
#6486
opened Jul 21, 2026 by
Solaris-star
Loading…
4 of 8 tasks
[DistillationTrainer refactor] Loss consumes
completion_mask
#6484
opened Jul 21, 2026 by
qgallouedec
Member
Loading…
[DistillationTrainer refactor] Emit
completion_mask alongside labels
#6482
opened Jul 21, 2026 by
qgallouedec
Member
Loading…
[DistillationTrainer refactor] Pin signature columns to
["prompt", "image", "images"]
#6481
opened Jul 21, 2026 by
qgallouedec
Member
Loading…
[DistillationTrainer refactor] Remove messages-format support and prompt-length config
#6480
opened Jul 21, 2026 by
qgallouedec
Member
Loading…
[DistillationTrainer refactor] Switch tests, docs, and example to prompt-only datasets
#6479
opened Jul 21, 2026 by
qgallouedec
Member
Loading…
[DistillationTrainer refactor] Fix
num_items_in_batch to count generated completion tokens ⚠️ changes loss values
#6478
opened Jul 21, 2026 by
qgallouedec
Member
Loading…
docs: complete GRPOConfig loss_type help
#6477
opened Jul 20, 2026 by
gowtham-sai-yadav
Loading…
3 of 8 tasks
fix(dpo): apply f_divergence_type consistently in apo_down loss
#6475
opened Jul 20, 2026 by
Ewertonslv
•
Draft
4 of 8 tasks
[DistillationTrainer refactor] Deprecate
messages-format datasets
#6474
opened Jul 20, 2026 by
qgallouedec
Member
Loading…
fix(dpo): keep apo_down on raw log-ratios for both terms
#6469
opened Jul 20, 2026 by
Solaris-star
Loading…
4 of 8 tasks
Raise a clear error for QLoRA +
vllm_mode="server" instead of a cryptic vLLM crash
#6440
opened Jul 18, 2026 by
DaoyuanLi2816
Contributor
Loading…
3 of 4 tasks
[GOLD] Support iterable datasets and collate each generation batch once per accumulation window
#6438
opened Jul 17, 2026 by
Strongich
Contributor
Loading…
Add privileged-context distillation to GOLD
#6437
opened Jul 17, 2026 by
eshwanthkartitr
Loading…
6 of 8 tasks
[OpenReward] Return None for unrewarded rollouts to prevent default-0.0 advantage poisoning
#6430
opened Jul 17, 2026 by
AUTHENSOR
Loading…
4 of 8 tasks
[SDPO] Port unscorable-mask reward defense from GRPO
#6429
opened Jul 17, 2026 by
AUTHENSOR
Loading…
5 of 8 tasks
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.