LLM4Decompile
Hanzhuo Tan, Qi Luo, Jing Li, et al.
LLM4Decompile is the first open-source LLM series trained to decompile binary code, outperforming GPT-4o and Ghidra by over 100% in re-executability.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Hanzhuo Tan, Qi Luo, Jing Li, et al.
LLM4Decompile is the first open-source LLM series trained to decompile binary code, outperforming GPT-4o and Ghidra by over 100% in re-executability.
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, et al.
Code Llama is a family of open-source LLMs for code based on Llama 2, achieving state-of-the-art performance in code generation, infilling, and long-context tasks.
Vincent Le Guilloux, Peter Schmidtke, Pierre Tufféry
Fpocket is an open-source platform for protein pocket detection using Voronoi tessellation and alpha spheres, achieving high accuracy and speed.
Unknown
olmOCR is an open-source Python toolkit that converts PDFs into linearized plain text while preserving structured content using a fine-tuned 7B VLM.
Unknown
Presents a detailed recipe for online iterative RLHF using fully open-source datasets, achieving state-of-the-art performance on conversation and instruction-following benchmarks.
Unknown
OmniMath introduces a comprehensive Olympiad-level math benchmark with 4428 problems across 33 sub-domains and 10 difficulty levels, using GPT-4o and an open-source verifier OmniJudge for rigorous evaluation.
Seungone Kim, Juyoung Suk, Shayne Longpre, et al.
Introduces Prometheus, a 13B fully open-source evaluation LLM trained on GPT-4-curated feedback data.
Unknown
A 137M parameter open-source English text embedding model with 8192 context length outperforming OpenAI on short and long-context tasks.
Liang Wang, Nan Yang, Xiaolong Huang, et al.
E5-Mistral-7B uses synthetic data from proprietary LLMs to fine-tune open-source decoder-only models for state-of-the-art text embeddings.
Unknown
Jina Embeddings v2 is an open-source model that extends BERT to handle up to 8192 tokens for long-document embeddings.
Unknown
Maya is an open-source multilingual multimodal model for vision-language understanding in eight languages, built on a toxicity-filtered dataset and fine-tuned on PALO 150K.
Unknown
BLIP-3 (xGen-MM) presents a comprehensive open-source framework for building large multimodal models with strong in-context learning.