Recurrent Memory Finds What LLMs Miss
Yuri Kuratov, A. Bulatov, Petr Anokhin, et al.
Recurrent memory augmentation enables GPT-2 to process sequences up to 11 million elements, far exceeding standard methods limited to 10,000 elements.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Yuri Kuratov, A. Bulatov, Petr Anokhin, et al.
Recurrent memory augmentation enables GPT-2 to process sequences up to 11 million elements, far exceeding standard methods limited to 10,000 elements.
P. Colombo, T. Pires, Malik Boudiaf, et al.
SaulLM-7B is a 7-billion-parameter LLM tailored for the legal domain, trained on over 30 billion tokens of English legal text and instruction-tuned to achieve state-of-the-art legal comprehension and generation.
Hanzhuo Tan, Qi Luo, Jing Li, et al.
LLM4Decompile is the first open-source LLM series trained to decompile binary code, outperforming GPT-4o and Ghidra by over 100% in re-executability.
Shubham Vatsal, Harsh Dubey
A survey of 44 papers on 39 prompt engineering methods across 29 NLP tasks, showing how structured prompts improve LLM performance without retraining.
Yuqing Yang, Yan Ma, Pengfei Liu
Proposes a progressive weak-to-strong reasoning framework where a strong model refines its own training data without human or advanced model input, significantly improving reasoning on GSM8K and MATH.
Tianyu Gu, Kang Liu, Brendan Dolan-Gavitt, et al.
This paper demonstrates that outsourced training of deep neural networks introduces security risks where adversaries can create backdoored networks that perform well on normal inputs but fail on attacker-chosen inputs.
Lavender Yao Jiang, Xujin Chris Liu, Nima Pour Nejatian, et al.
NYUTron, a large language model trained on unstructured clinical notes, serves as an all-purpose predictive engine for diverse clinical tasks, outperforming traditional models by 5-15% AUC.
Evan Shelhamer, Jonathan Long, Trevor Darrell
Fully convolutional networks trained end-to-end for semantic segmentation achieve state-of-the-art results by adapting classification nets and using a skip architecture for detailed predictions.
Kazuki Nonoyama, Ziang Liu, Tomofumi Fujiwara, et al.
This paper optimizes energy-efficient motion planning for dual-arm industrial robots by fine-tuning PID controllers with Genetic Algorithms and Particle Swarm Optimization.
Yuren Mao, Yuhang Ge, Yijiang Fan, et al.
A comprehensive survey categorizing LoRA variants for LLMs into downstream adaptation, cross-task generalization, efficiency, privacy, and applications.
DM Anisuzzaman, Jeffrey G. Malins, Paul A. Friedman, et al.
This review outlines major methodological approaches and steps for fine-tuning LLMs for specialized use cases, with examples from medical subspecialties.
Caleb Ziems, William A. Held, Omar Ahmed Shaikh, et al.
This paper provides a roadmap for using zero-shot LLMs as tools in computational social science, showing they can augment human annotation and bootstrapping creative generation tasks.