LearnLM
Unknown
LearnLM enhances Gemini for pedagogical instruction by co-training SFT and RLHF, outperforming leading LLMs in tutoring scenarios.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
LearnLM enhances Gemini for pedagogical instruction by co-training SFT and RLHF, outperforming leading LLMs in tutoring scenarios.
Unknown
CriticGPT uses RLHF to train a GPT-4-based model that critiques ChatGPT code outputs, helping humans catch bugs more accurately.
Unknown
Constrained Generative Policy Optimization with Mixture of Judges uses rule-based and LLM judges to constrain RLHF, improving multi-objective optimization and Pareto frontier.
Unknown
Presents a detailed recipe for online iterative RLHF using fully open-source datasets, achieving state-of-the-art performance on conversation and instruction-following benchmarks.
Unknown
RAFT iteratively fine-tunes generative models on top-ranked samples to align them with a reward function, improving stability and efficiency over RLHF.
Unknown
A benchmark dataset and code-base designed to evaluate reward models used in RLHF.
Unknown
LLaMa fine-tuned on only 1,000 curated examples matches or exceeds state-of-the-art RLHF models.