Geometric-Mean Policy Optimization
Yuzhong Zhao, Yue Liu, Junpeng Liu, et al.
GMPO improves GRPO stability by replacing arithmetic mean with geometric mean of token rewards, reducing outlier sensitivity and boosting reasoning performance.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Yuzhong Zhao, Yue Liu, Junpeng Liu, et al.
GMPO improves GRPO stability by replacing arithmetic mean with geometric mean of token rewards, reducing outlier sensitivity and boosting reasoning performance.
Zhi‐Hua Zhou
A comprehensive textbook covering ensemble methods including Boosting, Bagging, Random Forest, and diversity measures.
Robi Polikar
This paper reviews ensemble-based systems in decision making, covering conditions for their superiority over single classifiers, algorithms like bagging and boosting, combination rules, and future applications.
Alexey Natekin, Alois Knoll
This tutorial provides a comprehensive introduction to gradient boosting machines, covering methodology, model complexity, and practical applications.
Jerome H. Friedman
This paper introduces gradient boosting machines, a general paradigm for function estimation via stagewise additive expansions and steepest-descent minimization in function space.
Unknown
OmniParser converts UI screenshots into structured elements, boosting GPT-4V's ability to interact accurately with interfaces.
Unknown
V-STaR iteratively improves LLM reasoning by training a DPO verifier on correct and incorrect solutions, boosting generator and verifier performance.
Yiming Jia, Jiachen Li, Xiang Yue, et al.
MAmmoTH-VL 2 introduces VisualWebInstruct, a method using web search and LLMs to create a large multimodal instruction dataset, boosting VLM reasoning performance.
Unknown
MAmmoTH2 introduces a scalable pipeline to harvest 10 million natural instruction-response pairs from web corpora, boosting LLM reasoning without manual annotation.
Laila Rasmy, Yang Xiang, Ziqian Xie, et al.
Med-BERT adapts BERT to structured EHR data via pretraining on 28.5M patients, boosting disease prediction AUC by 1.21-6.14% and enabling small-dataset models to match those trained on tenfold larger data.