Kernel methods in machine learning
Thomas Hofmann, Bernhard Schölkopf, Alexander J. Smola
A comprehensive review of machine learning methods using positive definite kernels, formulated in reproducing kernel Hilbert spaces.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Thomas Hofmann, Bernhard Schölkopf, Alexander J. Smola
A comprehensive review of machine learning methods using positive definite kernels, formulated in reproducing kernel Hilbert spaces.
Arthur Gretton, Karsten Borgwardt, Malte J. Rasch, et al.
Proposes maximum mean discrepancy (MMD) in a reproducing kernel Hilbert space for two-sample testing, with distribution-free tests and quadratic-time computation.
Yan Hu, Qingyu Chen, Jingcheng Du, et al.
This paper shows that task-specific prompt engineering, incorporating medical knowledge and few-shot examples, significantly improves GPT-3.5 and GPT-4 performance on clinical NER tasks, though still below BioClinicalBERT.
Unknown
This paper proposes Structural LM, which extends BERT by incorporating 1D and 2D cell-level embeddings for structured text understanding.
Unknown
LAMBERT enhances RoBERTa with layout embeddings and relative bias for layout-aware document understanding without raw images.
Xinkun Huang, Jinyan Sang, Changrong Jin, et al.
Extends BERT with 2-D position and image embeddings for layout-aware document understanding.
Unknown
BEiT adapts BERT-style masked language modeling to vision by predicting discrete visual tokens for masked image patches.
Unknown
A suite of bilingual text embedding models supporting up to 8192 tokens, trained via modified BERT pre-training and multi-task contrastive learning.
Unknown
Jina Embeddings v2 is an open-source model that extends BERT to handle up to 8192 tokens for long-document embeddings.
Unknown
ColBERTv2 improves late-interaction retrieval by combining aggressive residual compression with denoised supervision to enhance quality and reduce storage.
Unknown
NeoBERT redefines bidirectional encoders with optimal depth-to-width ratio, 4K context, and 250M parameters, achieving state-of-the-art GLUE and MTEB scores.
Unknown
ModernBERT-Large-Instruct uses its MLM head for generative classification, achieving zero-shot performance rivaling larger LLMs.