[{"entity":{"id":"arxiv:1207.0580","text":"Improving neural networks by preventing co-adaptation of feature detectors. When a large feedforward neural network is trained on a small training set, it typically performs poorly on held-out test data. This \"overfitting\" is greatly reduced by randomly omitting half of the feature detectors on each training case. This prevents complex co-adaptations in which a feature detector is only helpful in the context of several other specific feature detectors. Instead, each neuron learns to detect a feature that is generally helpful for producing the correct answer given the combinatorially large variety of internal contexts in which it must operate. Random \"dropout\" gives big improvements on many benchmark tasks and sets new records for speech and object recognition.","hash":"ba809eb02ed6e3f20d589941315a31b0e6cc4cf3402827577916e76bf642d28e"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":0.4462922252986252,"interval":{"lower":-0.32751985352649005,"upper":1.2201043041237405,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":0.4462922252986252,"stddev":0.39480208103322206}},"context":{"rank":20,"percentile":0.1875,"z_score":-1.120009161658437,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:1301.3781","text":"Efficient Estimation of Word Representations in Vector Space. We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the previously best performing techniques based on different types of neural networks. We observe large improvements in accuracy at much lower computational cost, i.e. it takes less than a day to learn high quality word vectors from a 1.6 billion words data set. Furthermore, we show that these vectors provide state-of-the-art performance on our test set for measuring syntactic and semantic word similarities.","hash":"db1e7fae6176ef3f725bfd9261453b122f7294d64293b92f4eb2372337761327"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":0.7324422505311042,"interval":{"lower":-0.06555669232274575,"upper":1.5304411933849542,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":0.7324422505311042,"stddev":0.40714231778257653}},"context":{"rank":18,"percentile":0.2708333333333333,"z_score":-0.6347899460200335,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:1312.5602","text":"Playing Atari with Deep Reinforcement Learning. We present the first deep learning model to successfully learn control policies directly from high-dimensional sensory input using reinforcement learning. The model is a convolutional neural network, trained with a variant of Q-learning, whose input is raw pixels and whose output is a value function estimating future rewards. We apply our method to seven Atari 2600 games from the Arcade Learning Environment, with no adjustment of the architecture or learning algorithm. We find that it outperforms all previous approaches on six of the games and surpasses a human expert on three of them.","hash":"afd4ec4c41f4bf688c0c11fcece406c4b6290c549672871bebc22ffcb2865323"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.1329601592682739,"interval":{"lower":0.38167503618597143,"upper":1.8842452823505762,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.1329601592682739,"stddev":0.3833087362664808}},"context":{"rank":12,"percentile":0.5208333333333334,"z_score":0.04436073855074305,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:1312.6114","text":"Auto-Encoding Variational Bayes. How can we perform efficient inference and learning in directed probabilistic models, in the presence of continuous latent variables with intractable posterior distributions, and large datasets? We introduce a stochastic variational inference and learning algorithm that scales to large datasets and, under some mild differentiability conditions, even works in the intractable case. Our contributions are two-fold. First, we show that a reparameterization of the variational lower bound yields a lower bound estimator that can be straightforwardly optimized using standard stochastic gradient methods. Second, we show that for i.i.d. datasets with continuous latent variables per datapoint, posterior inference can be made especially efficient by fitting an approximate inference model (also called a recognition model) to the intractable posterior using the proposed lower bound estimator. Theoretical advantages are reflected in experimental results.","hash":"3e06a9f775bf6c383271b070228a8fdec51deffc43fd43310091e74c0cd1be3f"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":0.07826872779238281,"interval":{"lower":-0.7827293393612264,"upper":0.939266794945992,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":0.07826872779238281,"stddev":0.4392847281395965}},"context":{"rank":23,"percentile":0.0625,"z_score":-1.7440596842869156,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:1406.2661","text":"Generative Adversarial Networks. We propose a new framework for estimating generative models via an adversarial process, in which we simultaneously train two models: a generative model G that captures the data distribution, and a discriminative model D that estimates the probability that a sample came from the training data rather than G. The training procedure for G is to maximize the probability of D making a mistake. This framework corresponds to a minimax two-player game. In the space of arbitrary functions G and D, a unique solution exists, with G recovering the training data distribution and D equal to 1/2 everywhere. In the case where G and D are defined by multilayer perceptrons, the entire system can be trained with backpropagation. There is no need for any Markov chains or unrolled approximate inference networks during either training or generation of samples. Experiments demonstrate the potential of the framework through qualitative and quantitative evaluation of the generated samples.","hash":"d944af7424f005ec885627005229888301d192ea1306e28e75ff3e9b521c00c8"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":0.07914069680014779,"interval":{"lower":-0.732468896053458,"upper":0.8907502896537536,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":0.07914069680014779,"stddev":0.4140865269661254}},"context":{"rank":22,"percentile":0.10416666666666667,"z_score":-1.7425811028411688,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:1409.0473","text":"Neural Machine Translation by Jointly Learning to Align and Translate. Neural machine translation is a recently proposed approach to machine translation. Unlike the traditional statistical machine translation, the neural machine translation aims at building a single neural network that can be jointly tuned to maximize the translation performance. The models proposed recently for neural machine translation often belong to a family of encoder-decoders and consists of an encoder that encodes a source sentence into a fixed-length vector from which a decoder generates a translation. In this paper, we conjecture that the use of a fixed-length vector is a bottleneck in improving the performance of this basic encoder-decoder architecture, and propose to extend this by allowing a model to automatically (soft-)search for parts of a source sentence that are relevant to predicting a target word, without having to form these parts as a hard segment explicitly. With this new approach, we achieve a translation performance comparable to the existing state-of-the-art phrase-based system on the task of English-to-French translation. Furthermore, qualitative analysis reveals that the (soft-)alignments found by the model agree well with our intuition.","hash":"b6fe0a2067d9c7746d271dc6551d710dfff194b363f93d71656b3ac13a75dbd8"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":0.45793052816017976,"interval":{"lower":-0.42233743643061294,"upper":1.3381984927509725,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":0.45793052816017976,"stddev":0.44911630846469014}},"context":{"rank":19,"percentile":0.22916666666666666,"z_score":-1.100274310399005,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:1412.6980","text":"Adam: A Method for Stochastic Optimization. We introduce Adam, an algorithm for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimates of lower-order moments. The method is straightforward to implement, is computationally efficient, has little memory requirements, is invariant to diagonal rescaling of the gradients, and is well suited for problems that are large in terms of data and/or parameters. The method is also appropriate for non-stationary objectives and problems with very noisy and/or sparse gradients. The hyper-parameters have intuitive interpretations and typically require little tuning. Some connections to related algorithms, on which Adam was inspired, are discussed. We also analyze the theoretical convergence properties of the algorithm and provide a regret bound on the convergence rate that is comparable to the best known results under the online convex optimization framework. Empirical results demonstrate that Adam works well in practice and compares favorably to other stochastic optimization methods. Finally, we discuss AdaMax, a variant of Adam based on the infinity norm.","hash":"92e0cc684b5a00b57c801bdcd8e903b948593d3aef4ab2212fed36293895bbde"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":0.21305387918449248,"interval":{"lower":-0.5653902641127552,"upper":0.9914980224817401,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":0.21305387918449248,"stddev":0.39716537923328965}},"context":{"rank":21,"percentile":0.14583333333333334,"z_score":-1.5155070382228475,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:1502.03167","text":"Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. Training Deep Neural Networks is complicated by the fact that the distribution of each layer's inputs changes during training, as the parameters of the previous layers change. This slows down the training by requiring lower learning rates and careful parameter initialization, and makes it notoriously hard to train models with saturating nonlinearities. We refer to this phenomenon as internal covariate shift, and address the problem by normalizing layer inputs. Our method draws its strength from making normalization a part of the model architecture and performing the normalization for each training mini-batch. Batch Normalization allows us to use much higher learning rates and be less careful about initialization. It also acts as a regularizer, in some cases eliminating the need for Dropout. Applied to a state-of-the-art image classification model, Batch Normalization achieves the same accuracy with 14 times fewer training steps, and beats the original model by a significant margin. Using an ensemble of batch-normalized networks, we improve upon the best published result on ImageNet classification: reaching 4.9% top-5 validation error (and 4.8% test error), exceeding the accuracy of human raters.","hash":"bec96c9bfabb47b3cc0926fdd6be5bed165d5a85a00eaa0ab5002dbcf5933c51"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.5829182487546916,"interval":{"lower":0.7376279266857694,"upper":2.428208570823614,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.5829182487546916,"stddev":0.43127057248414397}},"context":{"rank":4,"percentile":0.8541666666666666,"z_score":0.8073462077058691,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:1512.03385","text":"Deep Residual Learning for Image Recognition. Deeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of learning unreferenced functions. We provide comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth. On the ImageNet dataset we evaluate residual nets with a depth of up to 152 layers---8x deeper than VGG nets but still having lower complexity. An ensemble of these residual nets achieves 3.57% error on the ImageNet test set. This result won the 1st place on the ILSVRC 2015 classification task. We also present analysis on CIFAR-10 with 100 and 1000 layers. The depth of representations is of central importance for many visual recognition tasks. Solely due to our extremely deep representations, we obtain a 28% relative improvement on the COCO object detection dataset. Deep residual nets are foundations of our submissions to ILSVRC & COCO 2015 competitions, where we also won the 1st places on the tasks of ImageNet detection, ImageNet localization, COCO detection, and COCO segmentation.","hash":"422d7212de8c479e9cb5f65eaa0f351adb730f7329908f427eaaac6dffc11007"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.519548512948966,"interval":{"lower":0.7055456731048594,"upper":2.333551352793073,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.519548512948966,"stddev":0.41530757134903407}},"context":{"rank":6,"percentile":0.7708333333333334,"z_score":0.699891338610939,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:1706.03762","text":"Attention Is All You Need. The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.","hash":"a7f589ed043c8f3224c66853294094e0f201ebf2bdda166ac53c0f9901a52d83"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.5764825666362412,"interval":{"lower":0.6467142672887009,"upper":2.5062508659837817,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.5764825666362412,"stddev":0.4743715812997655}},"context":{"rank":5,"percentile":0.8125,"z_score":0.7964333425852891,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:1810.04805","text":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers. As a result, the pre-trained BERT model can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks, such as question answering and language inference, without substantial task-specific architecture modifications. BERT is conceptually simple and empirically powerful. It obtains new state-of-the-art results on eleven natural language processing tasks, including pushing the GLUE score to 80.5% (7.7% point absolute improvement), MultiNLI accuracy to 86.7% (4.6% absolute improvement), SQuAD v1.1 question answering Test F1 to 93.2 (1.5 point absolute improvement) and SQuAD v2.0 Test F1 to 83.1 (5.1 point absolute improvement).","hash":"cebdf3e36baf3a68e4b85784321347f9df38928c9c8b86271f544191e1bab69e"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.6912227860723472,"interval":{"lower":0.7614544867248069,"upper":2.6209910854198877,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.6912227860723472,"stddev":0.4743715812997655}},"context":{"rank":3,"percentile":0.8958333333333334,"z_score":0.9909961745533586,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2001.08361","text":"Scaling Laws for Neural Language Models. We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude. Other architectural details such as network width or depth have minimal effects within a wide range. Simple equations govern the dependence of overfitting on model/dataset size and the dependence of training speed on model size. These relationships allow us to determine the optimal allocation of a fixed compute budget. Larger models are significantly more sample-efficient, such that optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.","hash":"9f1700feb48d4d375a4afd05b6400b320e17374641539e759d8278abd7118633"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":0.9357208054279329,"interval":{"lower":0.18515527379160102,"upper":1.686286337064265,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":0.9357208054279329,"stddev":0.38294159777363873}},"context":{"rank":16,"percentile":0.3541666666666667,"z_score":-0.2900943239140671,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2005.11401","text":"Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Large pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when fine-tuned on downstream NLP tasks. However, their ability to access and precisely manipulate knowledge is still limited, and hence on knowledge-intensive tasks, their performance lags behind task-specific architectures. Additionally, providing provenance for their decisions and updating their world knowledge remain open research problems. Pre-trained models with a differentiable access mechanism to explicit non-parametric memory can overcome this issue, but have so far been only investigated for extractive downstream tasks. We explore a general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation. We introduce RAG models where the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index of Wikipedia, accessed with a pre-trained neural retriever. We compare two RAG formulations, one which conditions on the same retrieved passages across the whole generated sequence, the other can use different passages per token. We fine-tune and evaluate our models on a wide range of knowledge-intensive NLP tasks and set the state-of-the-art on three open domain QA tasks, outperforming parametric seq2seq models and task-specific retrieve-and-extract architectures. For language generation tasks, we find that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline.","hash":"b2eea5be863ba95bff9130e5bb2e98123fadb49d7ce04bcf3896516fa435d3d0"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":0.926564501921751,"interval":{"lower":0.1135956392989802,"upper":1.739533364544522,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":0.926564501921751,"stddev":0.41478003195039326}},"context":{"rank":17,"percentile":0.3125,"z_score":-0.30562049555010745,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2005.14165","text":"Language Models are Few-Shot Learners. Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still requires task-specific fine-tuning datasets of thousands or tens of thousands of examples. By contrast, humans can generally perform a new language task from only a few examples or from simple instructions - something which current NLP systems still largely struggle to do. Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model. GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks, as well as several tasks that require on-the-fly reasoning or domain adaptation, such as unscrambling words, using a novel word in a sentence, or performing 3-digit arithmetic. At the same time, we also identify some datasets where GPT-3's few-shot learning still struggles, as well as some datasets where GPT-3 faces methodological issues related to training on large web corpora. Finally, we find that GPT-3 can generate samples of news articles which human evaluators have difficulty distinguishing from articles written by humans. We discuss broader societal impacts of this finding and of GPT-3 in general.","hash":"ef21bf51bd21637b8c41f07c2cbd84a27626212dde8fb45bba0fa05732c7bb5f"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.0338555996294263,"interval":{"lower":0.2738789724317561,"upper":1.7938322268270965,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.0338555996294263,"stddev":0.3877431771416685}},"context":{"rank":15,"percentile":0.3958333333333333,"z_score":-0.1236889991876442,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2006.11239","text":"Denoising Diffusion Probabilistic Models. We present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics. Our best results are obtained by training on a weighted variational bound designed according to a novel connection between diffusion probabilistic models and denoising score matching with Langevin dynamics, and our models naturally admit a progressive lossy decompression scheme that can be interpreted as a generalization of autoregressive decoding. On the unconditional CIFAR10 dataset, we obtain an Inception score of 9.46 and a state-of-the-art FID score of 3.17. On 256x256 LSUN, we obtain sample quality similar to ProgressiveGAN. Our implementation is available at https://github.com/hojonathanho/diffusion","hash":"14cc48ae181e0bc98484130f4649ab3f67db612d5544f9d6803992d26a88e6e9"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.4718746853675775,"interval":{"lower":0.6774439013975563,"upper":2.2663054693375986,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.4718746853675775,"stddev":0.4053218285561333}},"context":{"rank":9,"percentile":0.6458333333333334,"z_score":0.6190517258702853,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2101.03961","text":"Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. In deep learning, models typically reuse the same parameters for all inputs. Mixture of Experts (MoE) defies this and instead selects different parameters for each incoming example. The result is a sparsely-activated model -- with outrageous numbers of parameters -- but a constant computational cost. However, despite several notable successes of MoE, widespread adoption has been hindered by complexity, communication costs and training instability -- we address these with the Switch Transformer. We simplify the MoE routing algorithm and design intuitive improved models with reduced communication and computational costs. Our proposed training techniques help wrangle the instabilities and we show large sparse models may be trained, for the first time, with lower precision (bfloat16) formats. We design models based off T5-Base and T5-Large to obtain up to 7x increases in pre-training speed with the same computational resources. These improvements extend into multilingual settings where we measure gains over the mT5-Base version across all 101 languages. Finally, we advance the current scale of language models by pre-training up to trillion parameter models on the \"Colossal Clean Crawled Corpus\" and achieve a 4x speedup over the T5-XXL model.","hash":"e02527811d1c67f3a5d2303f571409e2801c258e71822a782594fd33f6a79b82"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.4719194832138844,"interval":{"lower":0.743776124728726,"upper":2.200062841699043,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.4719194832138844,"stddev":0.37150171351283595}},"context":{"rank":8,"percentile":0.6875,"z_score":0.6191276887356728,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2103.00020","text":"Learning Transferable Visual Models From Natural Language Supervision. State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is needed to specify any other visual concept. Learning directly from raw text about images is a promising alternative which leverages a much broader source of supervision. We demonstrate that the simple pre-training task of predicting which caption goes with which image is an efficient and scalable way to learn SOTA image representations from scratch on a dataset of 400 million (image, text) pairs collected from the internet. After pre-training, natural language is used to reference learned visual concepts (or describe new ones) enabling zero-shot transfer of the model to downstream tasks. We study the performance of this approach by benchmarking on over 30 different existing computer vision datasets, spanning tasks such as OCR, action recognition in videos, geo-localization, and many types of fine-grained object classification. The model transfers non-trivially to most tasks and is often competitive with a fully supervised baseline without the need for any dataset specific training. For instance, we match the accuracy of the original ResNet-50 on ImageNet zero-shot without needing to use any of the 1.28 million training examples it was trained on. We release our code and pre-trained model weights at https://github.com/OpenAI/CLIP.","hash":"d46b6ed8e5ee921266de2bd55ab5a216cfd2202001d869723ffd5d71f607f5bd"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.1675898604997528,"interval":{"lower":0.3527053551449646,"upper":1.982474365854541,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.1675898604997528,"stddev":0.41575740069121847}},"context":{"rank":11,"percentile":0.5625,"z_score":0.1030816715846614,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2106.09685","text":"LoRA: Low-Rank Adaptation of Large Language Models. An important paradigm of natural language processing consists of large-scale pre-training on general domain data and adaptation to particular tasks or domains. As we pre-train larger models, full fine-tuning, which retrains all model parameters, becomes less feasible. Using GPT-3 175B as an example -- deploying independent instances of fine-tuned models, each with 175B parameters, is prohibitively expensive. We propose Low-Rank Adaptation, or LoRA, which freezes the pre-trained model weights and injects trainable rank decomposition matrices into each layer of the Transformer architecture, greatly reducing the number of trainable parameters for downstream tasks. Compared to GPT-3 175B fine-tuned with Adam, LoRA can reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times. LoRA performs on-par or better than fine-tuning in model quality on RoBERTa, DeBERTa, GPT-2, and GPT-3, despite having fewer trainable parameters, a higher training throughput, and, unlike adapters, no additional inference latency. We also provide an empirical investigation into rank-deficiency in language model adaptation, which sheds light on the efficacy of LoRA. We release a package that facilitates the integration of LoRA with PyTorch models and provide our implementations and model checkpoints for RoBERTa, DeBERTa, and GPT-2 at https://github.com/microsoft/LoRA.","hash":"e13944f8deab595a4e56abdd3fe26685560cb879457cc80c403782098c913712"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.4895893681914814,"interval":{"lower":0.6621517292233113,"upper":2.3170270071596515,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.4895893681914814,"stddev":0.42216206069804596}},"context":{"rank":7,"percentile":0.7291666666666666,"z_score":0.6490901803422513,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2201.11903","text":"Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. We explore how generating a chain of thought -- a series of intermediate reasoning steps -- significantly improves the ability of large language models to perform complex reasoning. In particular, we show how such reasoning abilities emerge naturally in sufficiently large language models via a simple method called chain of thought prompting, where a few chain of thought demonstrations are provided as exemplars in prompting. Experiments on three large language models show that chain of thought prompting improves performance on a range of arithmetic, commonsense, and symbolic reasoning tasks. The empirical gains can be striking. For instance, prompting a 540B-parameter language model with just eight chain of thought exemplars achieves state of the art accuracy on the GSM8K benchmark of math word problems, surpassing even finetuned GPT-3 with a verifier.","hash":"d907466a828b3b0c1216d09c9c07479bfc52e21e67f849f4a8079bee31f7dab4"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.0806381322482794,"interval":{"lower":0.22471348470129326,"upper":1.9365627797952656,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.0806381322482794,"stddev":0.43669624874846236}},"context":{"rank":13,"percentile":0.4791666666666667,"z_score":-0.04436073855074305,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2203.02155","text":"Training language models to follow instructions with human feedback. Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users. In this paper, we show an avenue for aligning language models with user intent on a wide range of tasks by fine-tuning with human feedback. Starting with a set of labeler-written prompts and prompts submitted through the OpenAI API, we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback. We call the resulting models InstructGPT. In human evaluations on our prompt distribution, outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. Moreover, InstructGPT models show improvements in truthfulness and reductions in toxic output generation while having minimal performance regressions on public NLP datasets. Even though InstructGPT still makes simple mistakes, our results show that fine-tuning with human feedback is a promising direction for aligning language models with human intent.","hash":"708acb4f6d9c163d28b75339e49e4c45fa9a4798a6602c73012c529f03885aac"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.0766322449678862,"interval":{"lower":0.29809516865534025,"upper":1.855169321280432,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.0766322449678862,"stddev":0.39721279403701326}},"context":{"rank":14,"percentile":0.4375,"z_score":-0.051153446266037086,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2203.15556","text":"Training Compute-Optimal Large Language Models. We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\\times$ more more data. Chinchilla uniformly and significantly outperforms Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, greater than a 7% improvement over Gopher.","hash":"48591cf76eb0784af50c5f440507129a5b34e047c1bcc18bc716e4a5a17df903"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.8896122424421022,"interval":{"lower":1.0681635640733562,"upper":2.7110609208108483,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.8896122424421022,"stddev":0.4191064685554827}},"context":{"rank":2,"percentile":0.9375,"z_score":1.3274014442452087,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2205.14135","text":"FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. Transformers are slow and memory-hungry on long sequences, since the time and memory complexity of self-attention are quadratic in sequence length. Approximate attention methods have attempted to address this problem by trading off model quality to reduce the compute complexity, but often do not achieve wall-clock speedup. We argue that a missing principle is making attention algorithms IO-aware -- accounting for reads and writes between levels of GPU memory. We propose FlashAttention, an IO-aware exact attention algorithm that uses tiling to reduce the number of memory reads/writes between GPU high bandwidth memory (HBM) and GPU on-chip SRAM. We analyze the IO complexity of FlashAttention, showing that it requires fewer HBM accesses than standard attention, and is optimal for a range of SRAM sizes. We also extend FlashAttention to block-sparse attention, yielding an approximate attention algorithm that is faster than any existing approximate attention method. FlashAttention trains Transformers faster than existing baselines: 15% end-to-end wall-clock speedup on BERT-large (seq. length 512) compared to the MLPerf 1.1 training speed record, 3$\\times$ speedup on GPT-2 (seq. length 1K), and 2.4$\\times$ speedup on long-range arena (seq. length 1K-4K). FlashAttention and block-sparse FlashAttention enable longer context in Transformers, yielding higher quality models (0.7 better perplexity on GPT-2 and 6.4 points of lift on long-document classification) and entirely new capabilities: the first Transformers to achieve better-than-chance performance on the Path-X challenge (seq. length 16K, 61.4% accuracy) and Path-256 (seq. length 64K, 63.1% accuracy).","hash":"9a2f1476523d2f714687ee3d085e79fe16b63f728e8cffb5987d2d3f45179160"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":2.0445359323277428,"interval":{"lower":1.281842329372079,"upper":2.8072295352834065,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":2.0445359323277428,"stddev":0.38912938926309376}},"context":{"rank":1,"percentile":0.9791666666666666,"z_score":1.5901026312503304,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2212.08073","text":"Constitutional AI: Harmlessness from AI Feedback. As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and so we refer to the method as 'Constitutional AI'. The process involves both a supervised learning and a reinforcement learning phase. In the supervised phase we sample from an initial model, then generate self-critiques and revisions, and then finetune the original model on revised responses. In the RL phase, we sample from the finetuned model, use a model to evaluate which of the two samples is better, and then train a preference model from this dataset of AI preferences. We then train with RL using the preference model as the reward signal, i.e. we use 'RL from AI Feedback' (RLAIF). As a result we are able to train a harmless but non-evasive AI assistant that engages with harmful queries by explaining its objections to them. Both the SL and RL methods can leverage chain-of-thought style reasoning to improve the human-judged performance and transparency of AI decision making. These methods make it possible to control AI behavior more precisely and with far fewer human labels.","hash":"216d62fea9944299367be80aca03c7d9d1a62c4b2b99d7f389d4dae1b6e8317e"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":0.0,"interval":{"lower":-0.8480703269047678,"upper":0.8480703269047678,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":0.0,"stddev":0.4326889422983509}},"context":{"rank":24,"percentile":0.020833333333333332,"z_score":-1.8767784938609542,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}},{"entity":{"id":"arxiv:2312.00752","text":"Mamba: Linear-Time Sequence Modeling with Selective State Spaces. Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module. Many subquadratic-time architectures such as linear attention, gated convolution and recurrent models, and structured state space models (SSMs) have been developed to address Transformers' computational inefficiency on long sequences, but they have not performed as well as attention on important modalities such as language. We identify that a key weakness of such models is their inability to perform content-based reasoning, and make several improvements. First, simply letting the SSM parameters be functions of the input addresses their weakness with discrete modalities, allowing the model to selectively propagate or forget information along the sequence length dimension depending on the current token. Second, even though this change prevents the use of efficient convolutions, we design a hardware-aware parallel algorithm in recurrent mode. We integrate these selective SSMs into a simplified end-to-end neural network architecture without attention or even MLP blocks (Mamba). Mamba enjoys fast inference (5$\\times$ higher throughput than Transformers) and linear scaling in sequence length, and its performance improves on real data up to million-length sequences. As a general sequence model backbone, Mamba achieves state-of-the-art performance across several modalities such as language, audio, and genomics. On language modeling, our Mamba-3B model outperforms Transformers of the same size and matches Transformers twice its size, both in pretraining and downstream evaluation.","hash":"b19c4595e688ab07820ea3ff67ab832aa4bf45ac3201fc12ee16b3aec056f830"},"attribute":{"lens":"landmark-ml-abstracts","axis_key":"empirical-claim-density","axis_prompt":"the density of concrete, falsifiable empirical claims in the text: specific measured results, named benchmarks, and quantitative comparisons a reader could verify, relative to the amount of text","axis_prompt_hash":"6f55d5d419e15f88a1a3493b7553ebe95677ec0b7eedeed90ffb384cf681d0ba"},"estimate":{"point":1.4716770734973081,"interval":{"lower":0.6986919465105292,"upper":2.244662200484087,"coverage":0.95,"method":"normal_approximation"},"distribution":{"family":"normal_approximation","mean":1.4716770734973081,"stddev":0.3943801668299892}},"context":{"rank":10,"percentile":0.6041666666666666,"z_score":0.6187166391387185,"cohort_size":24,"scored_at":"2026-08-07T03:43:19.164000Z"},"provenance":{"run_id":"jrun_606e6cb6afd144b392c17b3be58bad5c","model":"google/gemini-3.1-flash-lite-preview","harness":"cardinal-harness","harness_version":"0.9.0","temperature":0.0,"seed":"4085052757057738998","comparison_budget":192,"comparisons_used":192,"stop_reason":"budget_exhausted","topk_error":9.34269818743402,"run_cost_nanodollars":49009000}}]