Competence-Gated Pooling of Language Models and Priors for Event Forecasting

Papers

arxiv:2609.12101

Published on Sep 10

· Submitted by

Aditi Tiwari on Sep 14

· University of Illinois at Urbana-Champaign

Upvote

2

Authors:

,

,

Abstract

A competence gate selectively integrates language model forecasts by estimating domain-level marginal value over external predictions, improving hybrid forecasting accuracy.

Generated by thinkingmachines/Inkling-Small

In hybrid forecasting, a language model is often one of several available signals. A system may already have a market, crowd, or statistical forecast and must decide whether the model adds useful information or should be ignored. The relevant target is therefore not standalone model accuracy, but relative competence, defined as the model's marginal value beyond the available external forecast. Under Brier loss, we characterize when model disagreement can improve an external forecast and derive the gain from using domain-specific rather than global pooling weights. We then introduce a competence gate that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast. Across 2,357 resolved binary questions and five language models, the gate improves the main external baseline from 0.0771 to 0.0732 Brier and significantly outperforms global forecast combinations. The gain remains significant under leakage controls and against a leakage-safe time-series prior on the pooled structured set, with separate evidence on FRED. In contrast, the gate gives no significant improvement on the official ForecastBench market subset, where it largely defers to the market. Across four Qwen models, verbal confidence does not reliably identify when the model outperforms the external forecast, while outcome-estimated competence supports better abstention decisions. These results provide a practical approach for selective model use based on measured marginal value.

View arXiv page View PDF Add to collection

Community

adititiwari19

Paper submitter 8 days ago

When should an LLM actually influence a forecast that already has a market, crowd, or statistical prediction?

Most forecasting work asks how accurate an LLM is in isolation. We instead study relative competence: the model’s marginal value beyond an already-available forecast.

We introduce competence-gated pooling, which learns domain-specific source weights from resolved outcomes and defers when the external source is stronger. Across 2,357 resolved forecasting questions, the gate improves Brier score from 0.0771 → 0.0732 and significantly outperforms global forecast combinations. On a strong live-market subset, however, it correctly finds no significant gain from adding the LLM.

We also find that self-reported LLM confidence does not reliably predict when the model beats the external forecast—while outcome-estimated competence enables substantially better selective prediction.

Takeaway: don’t ask only “Is the model good?” Ask “Does it add information beyond what we already know?”

librarian-bot

8 days ago

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on HF Mirror checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment

Upvote

2

Get this paper in your agent:

hf papers read 2609.12101

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.12101 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.12101 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.12101 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

← 返回资讯列表