Instruction Pretraining
Unknown
Instruction Pre-Training augments raw corpora with instruction-response pairs for supervised multitask pretraining of language models.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Instruction Pre-Training augments raw corpora with instruction-response pairs for supervised multitask pretraining of language models.
Unknown
Fine Web is a 15 trillion token dataset for pretraining LLMs that yields better-performing models than other open datasets.
Unknown
Dolma is an open corpus of three trillion tokens for language model pretraining, designed to support transparency and reproducibility in AI research.
Unknown
Introduces pretraining tasks on math reasoning and chart derendering to enhance chart comprehension in diverse visual language tasks.
Unknown
AIM shows that autoregressive pretraining on images scales effectively with model and data size, achieving strong performance on vision tasks.
Ze Liu, Han Hu, Yutong Lin, et al.
Swin Transformer V2 introduces post-norm, cosine attention, log-spaced position bias, and self-supervised pretraining to scale vision transformers to 3B parameters.
Unknown
SigLIP 2 improves multilingual vision-language encoders with captioning pretraining, self-supervised losses, and online data curation, offering native aspect ratio preservation.
Unknown
SimCLRv2 presents a semi-supervised learning framework combining unsupervised pretraining, supervised fine-tuning, and distillation with unlabeled data.
Unknown
Introduces a training method for developing ultra-long context LLMs with context windows extending up to 4 million tokens via efficient continued pretraining with YaRN-based scaling and instruction tuning.
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, et al.
Llemma is an open-source LLM for mathematics, built by continued pretraining of Code Llama on a curated dataset of scientific papers, web math, and code, enabling tool use and formal theorem proving.
Unknown
LLaMA 2 Long extends LLaMA 2 to handle up to 32,768 tokens via continual pretraining with modified RoPE and synthetic instruction data.
Zhilin Yang, Zihang Dai, Yiming Yang, et al.
XLNet combines autoregressive and autoencoding pretraining via permutation language modeling, outperforming BERT on NLP tasks.