← Library · ChineseDaily papers
Grouped by first public publication date.
2026-10-08
- One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails — Separates fail-open and fail-closed errors in open decision-model guardrails; misleading option names can reverse decisions when labels enter the model input.
- Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks — Compares Jev with 19 LLMs on 13 benchmarks; strong knowledge scores coexist with a marked weakness on mathematical word problems.
- Can Decision Models Understand Stance? Evaluating Jev Against General-Purpose LLMs — Jev matches GPT-5.6 on English VAST stance labels but trails stronger models on Chinese conversational stance, especially favor-versus-against distinctions.
- Can Jev be Your Q or Policy in Reinforcement Learning? — Uses frozen Jev as a policy reference, exploration judge and replay rater; gains depend on the learning role and supplied information.
- Adversarial Cues in Decision Models Used as Judges: The Role of Request Presentation — A colon edit increases false acceptance of explicitly wrong final answers under sorted request keys, while insertion-order requests reject both variants.
- TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models — Uses policy-driven generators to test calibration, framing, abstention and cost; good error ranking does not ensure calibrated or evidence-sensitive confidence.
- FastJEV: Understanding Redundancy for Compact JEV Inference — Combines context-state reuse, candidate-prefix sharing and layer pruning for OmniJev; compact execution does not consistently reduce measured latency.
- MetaEncoder: Exploring the Limit of Bi-Encoders for Multimodal System One Decision Making with Natural Language Interface — Trains a multimodal bi-encoder with request-to-candidate contrastive learning; including options in requests improves closed-set decisions while retaining cached candidate embeddings.
- Can a System-One LLM Perform Knowledge Tracing When Few or No Learners Are Logged? — Tests cold-start knowledge tracing with reader swaps and matched typed inputs; Jev benefits mainly from its model prior, while supervised methods catch up with more learners.
2026-10-07
2026-10-06
2026-10-05
2026-10-04
2026-10-03
2026-10-02
2026-10-01
2026-09-30
2026-09-29
2026-09-28
2026-09-27
2026-09-26
2026-09-25
2026-09-24
2026-09-23
2026-09-22
2026-09-21
2026-09-20
2026-09-19
2026-09-17