采样引擎

基本信息

属性 值
相关模块 me.lucko.spark.common.sampler
引擎选择 async(async-profiler,默认)或 java(内置引擎)
原生库 libasyncProfiler.so,从 jar 资源解压后加载
采样模式 EXECUTION(CPU 执行栈)/ ALLOCATION(内存分配)
配置文件 config.json 的 backgroundProfilerEngine(见 config.json)

spark 有两套采样引擎。async 模式调用原生 async-profiler;java 模式使用纯 JVM 实现。在 1.7.10 Forge 上两者都可用,但平台支持矩阵不同(见下)。

引擎选择

BackgroundSamplerManager.startSampler() 读取配置:

boolean forceJavaEngine = this.configuration.getString(OPTION_ENGINE, "async").equals("java");

判定方式是字符串精确等于 "java",因此配置值必须是小写且完全匹配 —— 写 "Java" 或 "JAVA" 都会被当作 async(BackgroundSamplerManager.java:104)。

引擎最终通过 SamplerBuilder.forceJavaSampler(forceJavaEngine) 传入。命令行侧对应 /spark profiler --force-java-sampler。

两种采样模式

SamplerMode 枚举定义了两个模式,差异不只是「采什么」,默认采样间隔也完全不同:

模式 采样对象 默认间隔 值转换
EXECUTION CPU 执行栈 4 ms 微秒 → 毫秒(value / 1000d)
ALLOCATION 内存分配 524287 字节(512 KiB) 不转换

1.7.10 上 /spark profiler --alloc 切到分配模式;不带 --interval 且 --interval <= 0 时回退到上表的默认间隔(SamplerMode.defaultInterval())。

async-profiler 原生库支持矩阵

AsyncProfilerAccess 用 OS 名与架构名查表决定加载哪个资源(AsyncProfilerAccess.java:150-155):

os arch 资源路径
linux amd64 spark/linux/amd64/libasyncProfiler.so
linux amd64-musl spark/linux/amd64-musl/libasyncProfiler.so
linux aarch64 spark/linux/aarch64/libasyncProfiler.so
macosx amd64 spark/macos/libasyncProfiler.so
macosx aarch64 spark/macos/libasyncProfiler.so

这 4 个 .so 是 spark-common/src/main/resources 下仅有的 4 个资源文件,已逐一在仓库中确认存在。

Windows 完全不在表内 —— 命中 libPath == null 时抛 UnsupportedSystemException(AsyncProfilerAccess.java:157-160)。因此 1.7.10 Forge 客户端在 Windows 上必须走 --force-java-sampler 或 backgroundProfilerEngine: "java"。

macOS 的 aarch64(Apple Silicon)与 amd64 共用同一个 macos 条目,即二者加载的是同一份二进制。

加载流程

  1. 由 jar 资源路径拼出 resource = "spark/" + libPath + "/libasyncProfiler.so"
  2. getClassLoader().getResource(resource),为 null 则抛 IllegalStateException("Could not find ... in spark jar file") —— 说明打包漏了对应平台的原生库
  3. 解压到 platform.getTemporaryFiles().create("spark-", "-libasyncProfiler.so.tmp")
  4. AsyncProfiler.getInstance(extractPath.toAbsolutePath().toString())
  5. UnsatisfiedLinkError 被包装为 NativeLoadingException 抛出

解压目标位于 spark 目录下的 tmp/ 子目录(TemporaryFiles,在 SparkPlatform 构造时绑定到 plugin.getPluginDirectory().resolve("tmp"))。该目录在 SparkPlatform.disable() 中被 temporaryFiles.deleteTemporaryFiles() 清理。

失败回退

后台采样器启动失败不会让 spark 崩溃,而是自动降级到 Java 引擎并持久化。BackgroundSamplerManager.initialise() 的逻辑(BackgroundSamplerManager.java:63-72):

  1. 检查 config.json 中的 _marker_background_profiler_failed 标记
  2. 若为 true:打印两条 WARNING —— 「It seems the background profiler failed to start when spark was last enabled. Sorry about that!」与「In the future, spark will try to use the built-in Java profiling engine instead.」
  3. 删除该标记,把 backgroundProfilerEngine 写成 "java" 并 save() 落盘
  4. 本次启动即以 Java 引擎继续

注意 _marker_background_profiler_failed 是内部状态键(常量 MARKER_FAILED),不是面向用户的配置项,仅在下划线前缀下由 spark 自行写入与清除。

数值

数值名 值 来源
EXECUTION 默认间隔 4 ms SamplerMode
ALLOCATION 默认间隔 524287 字节 SamplerMode
async 支持的 OS 组合 5 组(见上表) AsyncProfilerAccess
失败标记键 _marker_background_profiler_failed BackgroundSamplerManager.MARKER_FAILED

相关条目