🍊 The Latent Field

标签: moe

此标签下有11条笔记。

  • 2026年8月31日

    GShard

    • source
    • paper
    • moe
    • distributed-training
    • automatic-sharding
    • xla
  • 2026年8月31日

    Switch Transformers

    • source
    • paper
    • moe
    • sparse-model
    • expert-parallelism
    • mixed-precision
    • scaling
    • knowledge-distillation
  • 2026年8月31日

    DeepSeek-V2

    • source
    • paper
    • deepseek
    • moe
    • mla
    • kv-cache
    • long-context
    • distributed-training
  • 2026年8月31日

    DeepSeek-V3 Technical Report

    • source
    • paper
    • deepseek
    • moe
    • mla
    • fp8
    • multi-token-prediction
    • pretraining
  • 2026年6月30日

    Qwen3-Coder-Next Technical Report

    • source
    • paper
    • qwen
    • code-agent
    • mid-training
    • reinforcement-learning
    • tool-use
    • moe
  • 2026年6月01日

    Outrageously Large Neural Networks

    • source
    • paper
    • moe
    • sparse-model
  • 2026年6月01日

    DeepSeekMoE

    • source
    • paper
    • moe
    • deepseek
  • 2026年5月28日

    Meta Llama 4 Multimodal Intelligence

    • source
    • blog
    • llama
    • multimodal
    • moe
  • 2026年5月28日

    DeepSeek V4 Technical Documentation

    • source
    • report
    • deepseek
    • moe
    • long-context
    • agent
  • 2026年2月14日

    Mixture of Experts

    • architecture
    • scaling
    • moe
    • sparse-model
  • 2026年2月08日

    DeepSeek

    • model-family
    • deepseek
    • moe
    • reasoning-model

The Latent Field · An AI knowledge atlas built with Quartz © 2026