Papers
Communities
Events
Blog
Pricing
Search
Open menu
Home
Papers
2404.04475
Cited By
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
6 April 2024
Yann Dubois
Balázs Galambosi
Percy Liang
Tatsunori Hashimoto
ALM
Re-assign community
ArXiv
PDF
HTML
Papers citing
"Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators"
6 / 256 papers shown
Title
Learn Your Reference Model for Real Good Alignment
Alexey Gorbatovski
Boris Shaposhnikov
Alexey Malakhov
Nikita Surnachev
Yaroslav Aksenov
Ian Maksimov
Nikita Balagansky
Daniil Gavrilov
OffRL
45
25
0
15 Apr 2024
ODIN: Disentangled Reward Mitigates Hacking in RLHF
Lichang Chen
Chen Zhu
Davit Soselia
Jiuhai Chen
Tianyi Zhou
Tom Goldstein
Heng-Chiao Huang
M. Shoeybi
Bryan Catanzaro
AAML
42
51
0
11 Feb 2024
Aligner: Efficient Alignment by Learning to Correct
Jiaming Ji
Boyuan Chen
Hantao Lou
Donghai Hong
Borong Zhang
Xuehai Pan
Juntao Dai
Tianyi Qiu
Yaodong Yang
21
6
0
04 Feb 2024
SelectLLM: Can LLMs Select Important Instructions to Annotate?
Long Lei
Jaehyung Kim
Yueming Jin
Dongyeop Kang
SyDa
37
10
0
29 Jan 2024
Instruction Tuning for Large Language Models: A Survey
Shengyu Zhang
Linfeng Dong
Xiaoya Li
Sen Zhang
Xiaofei Sun
...
Jiwei Li
Runyi Hu
Tianwei Zhang
Fei Wu
Guoyin Wang
LM&MA
11
524
0
21 Aug 2023
Training language models to follow instructions with human feedback
Long Ouyang
Jeff Wu
Xu Jiang
Diogo Almeida
Carroll L. Wainwright
...
Amanda Askell
Peter Welinder
Paul Christiano
Jan Leike
Ryan J. Lowe
OSLM
ALM
301
11,730
0
04 Mar 2022
Previous
1
2
3
4
5
6