🍊 Latent Atlas 🍉

标签: evaluation

此标签下有10条笔记。

  • 2026年6月30日

    APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training

    • source
    • paper
    • agent
    • benchmark
    • evaluation
    • pretraining
  • 2026年6月30日

    Agentic Mid-training Strategy for General-Purpose Agent Capability Injection

    • source
    • report
    • agent
    • mid-training
    • agent-training
    • evaluation
  • 2026年6月01日

    KLong: Training LLM Agent for Extremely Long-horizon Tasks

    • source
    • paper
    • agents
    • long-horizon
    • reinforcement-learning
    • sft
    • evaluation
  • 2026年5月24日

    Hallucination Evaluation

    • evaluation
    • hallucination
  • 2026年5月24日

    Human Evaluation

    • evaluation
    • human-evaluation
  • 2026年5月24日

    LLM-as-a-Judge

    • evaluation
    • judge
  • 2026年5月24日

    Online Evaluation

    • evaluation
    • online-evaluation
  • 2026年5月23日

    Benchmark

    • evaluation
    • benchmark
  • 2026年5月23日

    Evaluation and Benchmark

    • application
    • evaluation
  • 2025年12月28日

    Perplexity

    • math
    • information-theory
    • evaluation

🍊 Latent Atlas 🍉 · An AI knowledge atlas built with Quartz © 2026