Code Llama
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, et al.
Code Llama is a family of open-source LLMs for code based on Llama 2, achieving state-of-the-art performance in code generation, infilling, and long-context tasks.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, et al.
Code Llama is a family of open-source LLMs for code based on Llama 2, achieving state-of-the-art performance in code generation, infilling, and long-context tasks.
Shen Nie, Fengqi Zhu, Zebin You, et al.
LLaDA is a diffusion model trained from scratch for language modeling that matches autoregressive LLMs like LLaMA3 8B across benchmarks and solves the reversal curse.
Unknown
An open-source family of heterogeneous reasoning models (Nano, Super, Ultra) with dynamic reasoning toggle, trained via NAS, distillation, and RL.
Unknown
Tulu V3 presents a fully open post-training recipe for Llama 3.1 models, achieving state-of-the-art performance via SFT, DPO, and RLVR.
Unknown
Meta's quantized Llama 3.2 and Llama 3.3 models reduce size and memory via QAT with LoRA and SpinQuant, enabling deployment on mobile devices.
Unknown
Introduces small and medium-sized vision LLMs (11B and 90B) alongside lightweight text-only models (1B and 3B).
Unknown
Hermes 3 introduces neutrally-aligned instruct and tool-use models fine-tuned from Llama 3.1, prioritizing strict adherence to user prompts without moral judgment.
Unknown
LLM Compiler is a suite of pre-trained models built on Code Llama for code optimization, including LLVM-IR assembly/disassembly and code size reduction.
Unknown
Llama 3.1 adds multimodal capabilities via compositional cross-attention adapters for vision and speech, achieving competitive performance with GPT-4V and strong video reasoning.
Unknown
Llama 3.1 introduces a family of multilingual language models up to 405B parameters, trained on 15T tokens, achieving performance comparable to GPT-4.
Unknown
Llama 3 introduces a family of 8B, 70B, and 405B parameter models trained on 15T tokens, achieving state-of-the-art performance across reasoning, coding, and multilingual tasks.
Unknown
H2O Danube 1.8B is a language model pre-trained on 1 trillion tokens using principles from LLaMA 2 and Mistral.