<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>基准测试 on Answer</title>
    <link>https://answer.freetools.me/tags/%E5%9F%BA%E5%87%86%E6%B5%8B%E8%AF%95/</link>
    <description>Recent content in 基准测试 on Answer</description>
    <generator>Hugo -- 0.152.2</generator>
    <language>zh-cn</language>
    <lastBuildDate>Fri, 13 Mar 2026 08:07:25 +0800</lastBuildDate>
    <atom:link href="https://answer.freetools.me/tags/%E5%9F%BA%E5%87%86%E6%B5%8B%E8%AF%95/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>大模型代码生成能力的边界与突破——从语法理解到语义推理的技术解析</title>
      <link>https://answer.freetools.me/%E5%A4%A7%E6%A8%A1%E5%9E%8B%E4%BB%A3%E7%A0%81%E7%94%9F%E6%88%90%E8%83%BD%E5%8A%9B%E7%9A%84%E8%BE%B9%E7%95%8C%E4%B8%8E%E7%AA%81%E7%A0%B4%E4%BB%8E%E8%AF%AD%E6%B3%95%E7%90%86%E8%A7%A3%E5%88%B0%E8%AF%AD%E4%B9%89%E6%8E%A8%E7%90%86%E7%9A%84%E6%8A%80%E6%9C%AF%E8%A7%A3%E6%9E%90/</link>
      <pubDate>Fri, 13 Mar 2026 08:07:25 +0800</pubDate>
      <guid>https://answer.freetools.me/%E5%A4%A7%E6%A8%A1%E5%9E%8B%E4%BB%A3%E7%A0%81%E7%94%9F%E6%88%90%E8%83%BD%E5%8A%9B%E7%9A%84%E8%BE%B9%E7%95%8C%E4%B8%8E%E7%AA%81%E7%A0%B4%E4%BB%8E%E8%AF%AD%E6%B3%95%E7%90%86%E8%A7%A3%E5%88%B0%E8%AF%AD%E4%B9%89%E6%8E%A8%E7%90%86%E7%9A%84%E6%8A%80%E6%9C%AF%E8%A7%A3%E6%9E%90/</guid>
      <description>深入分析大语言模型在代码生成任务中的真实能力边界，从语法理解、静态语义分析到动态语义推理三个层次展开，揭示模型幻觉问题、安全性隐患以及评估基准的局限性，帮助开发者正确理解和使用代码生成工具。</description>
    </item>
    <item>
      <title>大模型如何评估：从标准化考试到人类偏好的完整技术解析</title>
      <link>https://answer.freetools.me/%E5%A4%A7%E6%A8%A1%E5%9E%8B%E5%A6%82%E4%BD%95%E8%AF%84%E4%BC%B0%E4%BB%8E%E6%A0%87%E5%87%86%E5%8C%96%E8%80%83%E8%AF%95%E5%88%B0%E4%BA%BA%E7%B1%BB%E5%81%8F%E5%A5%BD%E7%9A%84%E5%AE%8C%E6%95%B4%E6%8A%80%E6%9C%AF%E8%A7%A3%E6%9E%90/</link>
      <pubDate>Wed, 11 Mar 2026 13:52:30 +0800</pubDate>
      <guid>https://answer.freetools.me/%E5%A4%A7%E6%A8%A1%E5%9E%8B%E5%A6%82%E4%BD%95%E8%AF%84%E4%BC%B0%E4%BB%8E%E6%A0%87%E5%87%86%E5%8C%96%E8%80%83%E8%AF%95%E5%88%B0%E4%BA%BA%E7%B1%BB%E5%81%8F%E5%A5%BD%E7%9A%84%E5%AE%8C%E6%95%B4%E6%8A%80%E6%9C%AF%E8%A7%A3%E6%9E%90/</guid>
      <description>深入解析大语言模型评估体系的演进历程。从MMLU、GSM8K等标准化基准测试，到Chatbot Arena的人类偏好排行，再到数据污染、基准饱和等核心挑战，全面揭示如何科学评估一个大模型的真正能力。</description>
    </item>
    <item>
      <title>负载测试为何总是测不准：从协调遗漏到统计陷阱的二十年反思</title>
      <link>https://answer.freetools.me/%E8%B4%9F%E8%BD%BD%E6%B5%8B%E8%AF%95%E4%B8%BA%E4%BD%95%E6%80%BB%E6%98%AF%E6%B5%8B%E4%B8%8D%E5%87%86%E4%BB%8E%E5%8D%8F%E8%B0%83%E9%81%97%E6%BC%8F%E5%88%B0%E7%BB%9F%E8%AE%A1%E9%99%B7%E9%98%B1%E7%9A%84%E4%BA%8C%E5%8D%81%E5%B9%B4%E5%8F%8D%E6%80%9D/</link>
      <pubDate>Mon, 09 Mar 2026 02:41:40 +0800</pubDate>
      <guid>https://answer.freetools.me/%E8%B4%9F%E8%BD%BD%E6%B5%8B%E8%AF%95%E4%B8%BA%E4%BD%95%E6%80%BB%E6%98%AF%E6%B5%8B%E4%B8%8D%E5%87%86%E4%BB%8E%E5%8D%8F%E8%B0%83%E9%81%97%E6%BC%8F%E5%88%B0%E7%BB%9F%E8%AE%A1%E9%99%B7%E9%98%B1%E7%9A%84%E4%BA%8C%E5%8D%81%E5%B9%B4%E5%8F%8D%E6%80%9D/</guid>
      <description>深入解析负载测试结果失准的根本原因。从Gil Tene提出的协调遗漏（Coordinated Omission）概念，到CMU关于开放/封闭系统模型的经典研究，再到Adrian Cockcroft对百分位数陷阱的分析，系统梳理二十年来性能测试领域的核心认知误区。涵盖YCSB工具设计缺陷的真实案例、延迟修正的数学原理，以及企业级性能测试的实践指南。</description>
    </item>
  </channel>
</rss>
