采集模型与快照缓存

[!INFO] Git Commit: 5ea0bc3 | Updated: 2026-10-01

本 mod 相对上游最大的架构改动(commit 5ea0bc3 perf(collectors): cache metrics, sample on server thread)是把 6 个需要读取世界状态的采集器从「抓取时现算」改为「在服务器线程上定时采样、 Prometheus 线程只读缓存快照」。

为什么需要缓存

Sampler 的 Javadoc(Sampler.java:12-21)给出了理由:

sample() runs only on the server thread (driven by CollectorScheduler), where touching Minecraft world state is safe. collect() runs on the HTTP thread and only returns the cached snapshot, so a scrape never races the tick loop nor blocks the server.

改造前,Prometheus 的 HTTP 线程会在 collect() 里直接遍历 world.loadedEntityList、 loadedTileEntityList 等列表——而这些列表正被服务器线程修改,构成数据竞争,且抓取会阻塞 tick。 改造后两个线程的职责完全分离:

方法 线程 职责
sample() 服务器线程 读取 Minecraft 世界状态,构造指标
collect() HTTP 线程 return this.cache.get(); —— 仅返回不可变快照

Sampler 抽象基类

Sampler extends io.prometheus.client.Collector implements Collector.Describable (Sampler.java:22)。两个构造参数都是 public final:

public final String name;         // 采集器名,用作自监控指标的 label
public final int intervalTicks;   // 刷新间隔,单位 tick

快照容器

private final AtomicReference<List<MetricFamilySamples>> cache = new AtomicReference<>(EMPTY);

Sampler.java:41。初值是 Collections.emptyList()。因为存的是 AtomicReference<List<MetricFamilySamples>> 而非可变对象,HTTP 线程读到的永远是某次 sample() 产出的完整列表,不会看到半构造状态。

刷新与失败处理

public void refreshAt(long nowMillis) {
    long start = System.nanoTime();
    List<MetricFamilySamples> samples;
    try {
        samples = this.sample();
    } catch (Exception e) {
        LOG.warn("Collector '{}' failed to refresh; keeping previous snapshot.", this.name, e);
        return;                       // ← 提前返回,不更新任何时间戳
    }
    this.lastRefreshSeconds = (System.nanoTime() - start) / 1e9;
    this.lastRefreshMillis = nowMillis;
    this.cache.set(samples != null ? samples : EMPTY);
}

Sampler.java:90-102。三点行为:

  1. 捕获 Exception 而非 Throwable——Error(如 OutOfMemoryError)会穿透到服务器线程;
  2. 失败时保留旧快照且不更新时间戳——lastRefreshMillis / lastRefreshSeconds 都不动, 因此 staleness_seconds 会持续增大,能被自监控指标捕捉到。代价是 refresh_duration_seconds 会一直显示上一次成功的耗时,而不是本次失败的耗时;
  3. sample() 返回 null 被规整为 EMPTY(:101)——Teams 采集器在 ServerUtilities 未就绪时正是返回 null(Teams.java:54),这条兜底让它不会 NPE。

refresh() 是 refreshAt(System.currentTimeMillis()) 的便捷包装(:79-81)。 refreshAt(long) 单独存在是为了可测试(SamplerTest 直接传固定毫秒值)。

到期判定

public boolean isDue(long tick) { return tick % this.intervalTicks == 0; }

Sampler.java:72-74。CollectorScheduler 持有单调递增的 tick 计数器,采样在 tick 100、200、 300… 时发生。intervalTicks 来自配置,范围下限为 1(见 配置),所以取模不会除零。

自监控读取器

public double lastRefreshSeconds()                        // 上次成功刷新的耗时(秒)
public double stalenessSecondsAt(long nowMillis)          // (now - lastRefreshMillis) / 1000.0
public double stalenessSeconds()                          // 上者包一层 System.currentTimeMillis()

Sampler.java:117-135。两者都是 volatile double / volatile long, 所以 SelfMetrics 可以在 HTTP 线程安全读取(SelfMetrics 的 Javadoc 明确说明了这一点)。

CollectorScheduler

CollectorScheduler 用 GTNHLib @EventBusSubscriber 订阅 TickEvent.ServerTickEvent:

@SubscribeEvent
public static void onServerTick(TickEvent.ServerTickEvent event) {
    if (Instance != null && event.phase == TickEvent.Phase.END) {
        Instance.tick();
    }
}

CollectorScheduler.java:79-84。只在 END 阶段推进,且用静态 Instance 作为门控—— Instance 为 null(导出器未运行)时完全跳过。Instance 在 initCollectors() 中被赋值 (PrometheusExporterMod.java:129),在 closeCollectors() 中被清空(:99)。

tick() 的实现(:60-67):

long t = ++this.tick;
for (Sampler sampler : this.samplers) {
    if (sampler.isDue(t)) sampler.refresh();
}

集合是 CopyOnWriteArrayList(:28),遍历期间即使有采集器被增删也不会抛 ConcurrentModificationException——/prometheus restart 时会清空并重建整个列表。

启动时预热

public void refreshAll() {
    for (Sampler sampler : this.samplers) sampler.refresh();
}

CollectorScheduler.java:73-77,在 initCollectors() 末尾被调用一次 (PrometheusExporterMod.java:151),注释说明用途是 Populate every snapshot once so the first scrape is not empty—— 忽略间隔立即采样所有采集器。该调用发生在 FMLServerStartedEvent 中,本身就在服务器线程上。

clear() 清空列表并把 tick 归零(:44-47)。

BaseCollector

仅 3 行有效代码(BaseCollector.java),作用是给 Sampler 补上服务器引用:

public abstract class BaseCollector extends Sampler {
    final MinecraftServer mc_server;                    // 包级可见
    protected BaseCollector(MinecraftServer mc_server, String name, int intervalTicks) {
        super(name, intervalTicks);
        this.mc_server = mc_server;
    }
}

它存在的价值是让 6 个 Minecraft 采集器不必各自重复一遍构造转发。

两种采集器形态

形态 基类 成员 刷新方式
采样型 BaseCollector Entities、TileEntities、Chunks、Players、PlayerStatistics、Teams CollectorScheduler 按间隔驱动 sample()
事件驱动型 直接 extends Collector Ticks 订阅 TickEvent.ServerTickEvent + TickEvent.WorldTickEvent
自监控型 直接 extends Collector SelfMetrics 不缓存,每次 collect() 现读 volatile 字段

Ticks 之所以不用采样模型:tick 耗时必须由 START / END 事件配对计时, 间隔采样无法测出单次 tick 的耗时。SelfMetrics 无需缓存,因为它只读 volatile 标量。

关闭流程

stopExporter() → closeCollectors()(PrometheusExporterMod.java:93-106)按序执行:

  1. scheduler.clear() 并置 scheduler = null;
  2. CollectorScheduler.Instance = null(停止响应 tick 事件);
  3. Ticks.Instance = null(停止 tick 计时);
  4. CollectorRegistry.defaultRegistry.clear()(注销全部采集器)。

之后 closeHttpServer() 关闭 HTTP 服务器(:111-119)。顺序很重要:先断开 tick 事件, 再清空注册表,避免采样过程中访问已被清空的采集器。源码注释(:112-114、:172-173) 强调了 close() 的必要性——否则第二次加载存档时端口被占用会导致客户端崩溃。