返回博客列表

Supervision:把模型输出变成能交付的业务功能

Roboflow 出品的计算机视觉「胶水层」。它不训练模型也不做推理,专门把各家模型吐出来的原始结果,变成越线计数、区域告警、轨迹追踪这些真能交付的东西

做视觉项目最烦的一件事:模型换一次,后处理代码就得重写一遍。YOLO 返回一种结构,RF-DETR 返回另一种,Transformers 又是第三种。

Supervision 干的就是这件事——把二十多家推理框架的输出洗成同一个数据结构,然后在上面配齐一整套业务工具:画框、追踪 ID、越线计数、区域人数、数据集格式互转、mAP 评估。

版本GitHub StarPyPI 月下载协议Python
0.30.138,000+100 万+MIT(商用无限制)≥ 3.10

下面是一份可以照着跑的完整教程,从装环境到「路口车流统计」这种可交付的成品。


1. 它到底是什么

一句话:Supervision 不训练模型、不做推理,它把「模型吐出来的原始结果」变成「能用的业务功能」。

它在整个流程中的位置

Supervision 在计算机视觉流程中的位置 上中下三层:各家推理框架输出格式互不相同,经 Supervision 洗成统一的 sv.Detections 结构并接上工具链,向下产出标注、追踪、计数告警等可交付的业务功能。 推理框架 · 输出格式各不相同 SUPERVISION · 胶水层 你的业务 · 直接可交付 洗成一种格式 业务代码不变 Ultralytics YOLOv5 · v8 · v11 RF-DETR Detectron2 · MMDet SAM · VLM 分割 · 大模型 · OCR sv.Detections 统一数据结构 · 换模型只改这一行 工具链 30+ Annotator · ByteTrack · DetectionsSmoother LineZone · PolygonZone · InferenceSlicer · Dataset · Metrics 标注与打码 画框 · 标签 · 马赛克 跨帧追踪 ID · 轨迹 · 热力图 计数与告警 越线 · 区域 · 精度评估

为什么需要它

没有 Supervision 时,你换一次模型(YOLO → RF-DETR)就得重写一遍后处理代码, 因为每家返回的结果结构都不一样。有了它:

detections = sv.Detections.from_ultralytics(results)   # 换模型
detections = sv.Detections.from_transformers(results)  # 只改这一行
# ↓ 后面所有代码完全不变

这就是它最大的价值:解耦。 模型可以随便换,业务代码一行不动。

项目现状(2026-09)

项值
最新版本0.30.1
GitHub Star38,000+
PyPI 月下载量100 万+
协议MIT(商用无限制)
Python≥ 3.10
维护方Roboflow(商业公司,持续维护)

2. 安装

pip install supervision

配套装一个推理框架(教程用 Ultralytics YOLO 举例):

pip install ultralytics

验证:

import supervision as sv
print(sv.__version__)

约定:社区习惯 import supervision as sv,本文全部沿用。


3. 核心概念:Detections

整个库只有一个核心数据结构 —— sv.Detections。理解它 = 掌握 80% 的库。

3.1 它长什么样

sv.Detections(
    xyxy,        # (n, 4) float  必填  每个框的 [x1, y1, x2, y2]
    mask,        # (n, H, W) bool 可选  分割掩码
    confidence,  # (n,) float    可选  置信度
    class_id,    # (n,) int      可选  类别编号
    tracker_id,  # (n,) int      可选  追踪 ID(追踪后才有)
    data,        # dict          可选  自定义字段,如 class_name
    metadata,    # dict          可选  集合级信息,如视频名、时间戳
)

关键理解:它本质是一堆 numpy 数组的容器,n 是检测到的目标个数。 所以它天然支持 numpy 式的过滤和切片。

3.2 从各家模型导入(from_* 方法)

推理来源方法
Ultralytics (YOLOv8/v11)sv.Detections.from_ultralytics(results)
Roboflow Inferencesv.Detections.from_inference(results)
HuggingFace Transformerssv.Detections.from_transformers(results)
RF-DETR直接返回 Detections,无需转换
YOLOv5sv.Detections.from_yolov5(results)
YOLO-NASsv.Detections.from_yolo_nas(results)
Detectron2sv.Detections.from_detectron2(results)
MMDetectionsv.Detections.from_mmdetection(results)
SAM / SAM3sv.Detections.from_sam() / from_sam3()
PaddleDetectionsv.Detections.from_paddledet(results)
TensorFlowsv.Detections.from_tensorflow(results)
DeepSparse / NCNNfrom_deepsparse() / from_ncnn()
EasyOCRsv.Detections.from_easyocr(results)
Azure AI Visionsv.Detections.from_azure_analyze_image(results)
大模型 VLM/LMMsv.Detections.from_vlm() / from_lmm()

3.3 过滤:像操作 DataFrame 一样

# 只保留置信度 > 0.5 的
detections = detections[detections.confidence > 0.5]

# 只保留 class_id == 0(比如"人")
detections = detections[detections.class_id == 0]

# 只保留框面积 > 1000 像素的
detections = detections[detections.box_area > 1000]

# 多条件组合(注意用 & 不是 and,用括号包起来)
detections = detections[
    (detections.confidence > 0.5) & (detections.class_id == 0)
]

# 按数量
print(len(detections))

3.4 常用方法

detections.with_nms(threshold=0.45)          # 非极大值抑制,去重叠框
detections.with_soft_nms(sigma=0.5)          # Soft-NMS
detections.with_nmm(threshold=0.5)           # 非极大值合并(融合而非丢弃)

detections.area                              # 掩码面积(有 mask 时)
detections.box_area                          # 框面积
detections.box_aspect_ratio                  # 宽高比

detections.get_anchors_coordinates(sv.Position.BOTTOM_CENTER)  # 取锚点坐标
detections.get_data("class_name")            # 取自定义字段
sv.Detections.merge([det_a, det_b])          # 合并多组检测
sv.Detections.empty()                        # 空检测(占位用)
detections.is_empty()                        # 判空

get_anchors_coordinates 很重要:越线/区域判定都是靠「用框上的哪个点代表这个物体」。 行人通常用 BOTTOM_CENTER(脚底),车辆有时用 CENTER。


4. 第一课:图片检测 + 标注

4.1 最小可运行例子

import cv2
import supervision as sv
from ultralytics import YOLO

# 1) 跑模型
model = YOLO("yolov8n.pt")
image = cv2.imread("input.jpg")
results = model(image)[0]

# 2) 转成 Supervision 的统一结构
detections = sv.Detections.from_ultralytics(results)

# 3) 标注
box_annotator = sv.BoxAnnotator()
label_annotator = sv.LabelAnnotator()

annotated = box_annotator.annotate(scene=image.copy(), detections=detections)
annotated = label_annotator.annotate(scene=annotated, detections=detections)

cv2.imwrite("output.jpg", annotated)

注意三个模式:

  • Annotator 是先创建对象、再反复调用 .annotate()(性能考虑,别在循环里 new)
  • .annotate() 返回图像,所以可以链式叠加
  • 传 image.copy() 而不是 image,避免原图被就地改写

4.2 自定义标签

默认标签是「类别名 置信度」。想自定义:

labels = [
    f"{class_name} {confidence:.2f}"
    for class_name, confidence
    in zip(detections["class_name"], detections.confidence)
]

annotated = label_annotator.annotate(
    scene=annotated, detections=detections, labels=labels
)

4.3 Annotator 全家桶(30+ 个)

按用途分五类,都可以叠加使用:

类别Annotator效果
轮廓类BoxAnnotator标准矩形框
RoundBoxAnnotator圆角框
BoxCornerAnnotator只画四个角(科技感)
OrientedBoxAnnotator旋转框(OBB,如遥感、文字)
CircleAnnotator / EllipseAnnotator圆 / 椭圆(体育赛事常用)
PolygonAnnotator多边形轮廓
填充类ColorAnnotator半透明色块填充
MaskAnnotator实例分割掩码(需要 detections.mask)
HaloAnnotator光晕效果
标记类DotAnnotator中心点圆点
TriangleAnnotator三角标记(指向目标)
文字类LabelAnnotator带背景的文字标签
RichLabelAnnotator自定义字体(中文标签必须用这个)
图像处理类BlurAnnotator模糊目标区域(打码,如人脸/车牌)
PixelateAnnotator马赛克
追踪/聚合类TraceAnnotator运动轨迹线(需要 tracker_id)
HeatMapAnnotator热力图(人流密度分析)
PercentageBarAnnotator进度条形式显示置信度
其它IconAnnotator贴自定义图标
BackgroundOverlayAnnotator背景变暗,突出目标
ComparisonAnnotator两组检测结果对比(调参神器)
CropAnnotator把目标裁剪放大显示在旁边

中文标签的坑:LabelAnnotator 底层用 OpenCV 画字,不支持中文,会显示成 ???。 中文必须用 RichLabelAnnotator 并指定字体文件:

label_annotator = sv.RichLabelAnnotator(font_path="C:/Windows/Fonts/msyh.ttc")

4.4 颜色控制

box_annotator = sv.BoxAnnotator(
    color=sv.ColorPalette.DEFAULT,
    color_lookup=sv.ColorLookup.CLASS,   # 按类别上色(默认)
    # sv.ColorLookup.TRACK  → 按追踪 ID 上色(每个目标一个固定颜色)
    # sv.ColorLookup.INDEX  → 按序号上色
    thickness=2,
)

做追踪演示时,把 color_lookup 设成 TRACK,同一个人全程一个颜色,观感立刻专业。


5. 第二课:视频处理与目标追踪

5.1 视频工具三件套

import supervision as sv

# 读元信息
video_info = sv.VideoInfo.from_video_path("input.mp4")
print(video_info.width, video_info.height, video_info.fps, video_info.total_frames)

# 逐帧生成器(省内存,不会一次读进来)
for frame in sv.get_video_frames_generator("input.mp4"):
    ...

# 只取部分帧
for frame in sv.get_video_frames_generator("input.mp4", start=100, end=300, stride=2):
    ...

5.2 sv.process_video:视频处理的标准范式

这是全库最实用的一个函数,帮你搞定「读帧 → 处理 → 写回」的全部样板代码:

import numpy as np
import supervision as sv
from ultralytics import YOLO

model = YOLO("yolov8n.pt")
box_annotator = sv.BoxAnnotator()

def callback(frame: np.ndarray, index: int) -> np.ndarray:
    results = model(frame)[0]
    detections = sv.Detections.from_ultralytics(results)
    return box_annotator.annotate(frame.copy(), detections=detections)

sv.process_video(
    source_path="input.mp4",
    target_path="output.mp4",
    callback=callback
)

范式要点:你只需要写 callback(帧, 帧号) -> 处理后的帧,其余全部交给它。 分辨率、帧率、编码器都自动继承源视频。

5.3 目标追踪:ByteTrack

检测只知道「这一帧有 3 个人」;追踪才知道「这 3 个人跟上一帧是不是同一批」。 Supervision 内置了 ByteTrack(不需要额外装库):

import numpy as np
import supervision as sv
from ultralytics import YOLO

model = YOLO("yolov8n.pt")
tracker = sv.ByteTrack()             # ← 追踪器,全程只创建一次
box_annotator = sv.BoxAnnotator()
label_annotator = sv.LabelAnnotator()

def callback(frame: np.ndarray, _: int) -> np.ndarray:
    results = model(frame)[0]
    detections = sv.Detections.from_ultralytics(results)
    detections = tracker.update_with_detections(detections)   # ← 关键这一行

    labels = [
        f"#{tracker_id} {class_name}"
        for class_name, tracker_id
        in zip(detections.data["class_name"], detections.tracker_id)
    ]

    annotated = box_annotator.annotate(frame.copy(), detections=detections)
    return label_annotator.annotate(annotated, detections=detections, labels=labels)

sv.process_video("people-walking.mp4", "result.mp4", callback=callback)

调用 update_with_detections() 后,detections.tracker_id 就被填上了。

必须注意:tracker = sv.ByteTrack() 一定要放在 callback 外面。 放里面等于每帧重建追踪器,ID 会一直从 1 开始跳,追踪完全失效 —— 这是最高频的新手错误。

5.4 加轨迹线

trace_annotator = sv.TraceAnnotator(trace_length=30)   # 保留最近 30 帧的轨迹

def callback(frame, _):
    results = model(frame)[0]
    detections = sv.Detections.from_ultralytics(results)
    detections = tracker.update_with_detections(detections)

    annotated = box_annotator.annotate(frame.copy(), detections=detections)
    return trace_annotator.annotate(annotated, detections=detections)

5.5 抖动平滑

框跳动厉害时,加一个平滑器:

smoother = sv.DetectionsSmoother()

detections = tracker.update_with_detections(detections)
detections = smoother.update_with_detections(detections)   # 追踪之后

注意顺序:先追踪、后平滑。平滑器依赖 tracker_id 来判断哪些框属于同一个目标。


6. 第三课:越线计数(LineZone)

典型场景:门口进出人数、路口车流量、传送带计件。

6.1 原理

画一条虚拟线,物体的锚点从线的一侧穿到另一侧就计一次数。 必须先有追踪 ID(不然没法判断”同一个物体”跨越了)。

LineZone 越线计数与 PolygonZone 区域计数的判定方式 左侧:目标的锚点从计数线一侧穿到另一侧时累加 in_count 或 out_count。右侧:只有锚点落在多边形内才计入 current_count,框角伸进区域但锚点在外的目标不计。 LINEZONE · 越线计数 POLYGONZONE · 区域计数 in out 计一次 不计入 实心圆点 = 锚点(行人默认取脚底中心) 锚点从一侧穿到另一侧 → 计一次 in_count / out_count 分方向累加 锚点落在区内才算 current_count 左上那个:框角伸进来了,锚点在外 → 不计入,这就是选错锚点的坑

6.2 API

sv.LineZone(
    start: sv.Point,
    end: sv.Point,
    triggering_anchors = (            # 用框上哪些点判定穿越
        sv.Position.TOP_LEFT,
        sv.Position.TOP_RIGHT,
        sv.Position.BOTTOM_LEFT,
        sv.Position.BOTTOM_RIGHT,
    ),
    minimum_crossing_threshold: int = 1,   # 需连续多少帧确认,抗抖动
)

关键属性:

属性含义
in_count从外向内穿越的累计数
out_count从内向外穿越的累计数
in_count_per_class分类别的进入计数(dict)
out_count_per_class分类别的离开计数(dict)

6.3 完整示例

import numpy as np
import supervision as sv
from ultralytics import YOLO

model = YOLO("yolov8n.pt")
tracker = sv.ByteTrack()

# 定义计数线(坐标是像素,按你的视频分辨率来定)
line_zone = sv.LineZone(
    start=sv.Point(x=0, y=500),
    end=sv.Point(x=1920, y=500),
)
line_annotator = sv.LineZoneAnnotator(
    thickness=2,
    text_thickness=2,
    text_scale=1.0,
    custom_in_text="进",     # 中文这里同样受 OpenCV 限制,英文更稳
    custom_out_text="出",
)
box_annotator = sv.BoxAnnotator()

def callback(frame: np.ndarray, _: int) -> np.ndarray:
    results = model(frame)[0]
    detections = sv.Detections.from_ultralytics(results)
    detections = detections[detections.class_id == 0]        # 只算"人"
    detections = tracker.update_with_detections(detections)

    line_zone.trigger(detections)                            # ← 更新计数

    annotated = box_annotator.annotate(frame.copy(), detections=detections)
    return line_annotator.annotate(annotated, line_counter=line_zone)

sv.process_video("input.mp4", "output.mp4", callback=callback)

print(f"进入: {line_zone.in_count}, 离开: {line_zone.out_count}")

trigger() 返回 (crossed_in, crossed_out) 两个布尔数组, 需要”谁穿过去了”的明细时接住它;只要总数的话直接调用不接返回值即可。


7. 第四课:区域计数(PolygonZone)

典型场景:某区域实时人数、车位占用、危险区域闯入告警。

7.1 API

sv.PolygonZone(
    polygon: np.ndarray,              # (N, 2) 的顶点坐标
    triggering_anchors = (sv.Position.BOTTOM_CENTER,),   # 默认脚底中心
    require_all_anchors: bool = True, # True=所有锚点都在内才算;False=任一即可
)
成员说明
trigger(detections)返回布尔数组,标记哪些检测在区域内
current_count当前区域内目标数
mask区域的二维布尔掩码

7.2 完整示例

import numpy as np
import supervision as sv
from ultralytics import YOLO

model = YOLO("yolov8n.pt")

polygon = np.array([[300, 200], [900, 200], [900, 800], [300, 800]])
zone = sv.PolygonZone(polygon=polygon)
zone_annotator = sv.PolygonZoneAnnotator(
    zone=zone,
    color=sv.Color.RED,
    thickness=4,
    text_scale=1.5,
    opacity=0.2,          # 区域半透明填充
)
box_annotator = sv.BoxAnnotator()

def callback(frame: np.ndarray, _: int) -> np.ndarray:
    results = model(frame)[0]
    detections = sv.Detections.from_ultralytics(results)

    mask = zone.trigger(detections=detections)    # ← 布尔数组
    detections = detections[mask]                 # 只保留区域内的

    annotated = box_annotator.annotate(frame.copy(), detections=detections)
    annotated = zone_annotator.annotate(scene=annotated)

    if zone.current_count > 10:
        print("告警:区域人数超限")
    return annotated

sv.process_video("input.mp4", "output.mp4", callback=callback)

triggering_anchors 怎么选:行人用 BOTTOM_CENTER(脚踩在哪算在哪), 空中/俯视目标用 CENTER。选错会导致”人明明在区域外但框的角伸进来了也被计数”。

7.3 多边形怎么画出来

手写坐标很痛苦,用 Roboflow 的在线工具:https://polygonzone.roboflow.com 上传一张截图,鼠标点几下,直接生成 numpy 数组代码。


8. 第五课:小目标检测(InferenceSlicer)

8.1 问题

4K 航拍图直接丢进 YOLO,模型内部会把图缩到 640×640, 原本 20 像素的小车缩成 3 像素 —— 检不出来。

8.2 方案:切片推理(SAHI)

把大图切成若干 640×640 的小块,分别推理,再把结果拼回原图坐标系并去重。

InferenceSlicer 切片推理的工作方式 大图被切成若干带重叠的小块分别送进模型,小目标在小块里相对变大因而能被检出,各块结果再映射回原图坐标并用 NMS 去掉重叠框。 切成 6 块 · 每块 640×640 · overlap 100 原图 4000×3000 小目标仅约 20 像素 映射回原图坐标 重叠部分 NMS 去重 代价:4K 图约切 30 块,推理耗时约 30 倍 —— 只适合离线大图,不适合实时视频

8.3 完整签名

sv.InferenceSlicer(
    callback,                                    # 必填:单块推理函数
    slice_wh = 640,                              # 切片尺寸
    overlap_wh = 100,                            # 切片重叠像素
    overlap_filter = OverlapFilter.NON_MAX_SUPPRESSION,  # 拼接去重策略
    iou_threshold = 0.5,
    overlap_metric = OverlapMetric.IOU,
    thread_workers = 1,                          # 并行线程数
    compact_masks = False,
    batch_size = 1,
)

8.4 用法

import cv2
import supervision as sv
from ultralytics import YOLO

model = YOLO("yolo11m.pt")

def callback(tile) -> sv.Detections:
    results = model(tile)[0]
    return sv.Detections.from_ultralytics(results)

slicer = sv.InferenceSlicer(
    callback=callback,
    slice_wh=640,
    overlap_wh=100,
    thread_workers=4,      # 有 GPU 显存余量时开并行提速
)

image = cv2.imread("aerial.jpg")
detections = slicer(image)     # 像调函数一样调用

调参经验:

  • 目标经常在切片边界被切断 → 加大 overlap_wh
  • 太慢 → 减小 overlap_wh 或加大 slice_wh
  • 代价:切片数量 = 面积/切片面积,4K 图约 30 块,推理耗时约 30 倍,不适合实时视频

9. 第六课:数据集管理与格式转换

sv.DetectionsDataset 把 YOLO / COCO / Pascal VOC 三种格式打通, 互转只要两行。

9.1 加载

import supervision as sv

# YOLO 格式
ds = sv.DetectionDataset.from_yolo(
    images_directory_path="dataset/images",
    annotations_directory_path="dataset/labels",
    data_yaml_path="dataset/data.yaml",
    is_obb=False,              # 旋转框数据集设 True
    show_progress=True,
)

# COCO 格式
ds = sv.DetectionDataset.from_coco(
    images_directory_path="dataset/images",
    annotations_path="dataset/_annotations.coco.json",
)

# Pascal VOC 格式
ds = sv.DetectionDataset.from_pascal_voc(
    images_directory_path="dataset/images",
    annotations_directory_path="dataset/annotations",
)

print(ds.classes)     # 类别列表
print(len(ds))        # 图片数

9.2 导出(= 格式转换)

# COCO → YOLO,就这两步
ds = sv.DetectionDataset.from_coco(...)
ds.as_yolo(
    images_directory_path="out/images",
    annotations_directory_path="out/labels",
    data_yaml_path="out/data.yaml",
)

# YOLO → COCO
ds.as_coco(
    images_directory_path="out/images",
    annotations_path="out/annotations.json",
)

# → Pascal VOC
ds.as_pascal_voc(
    images_directory_path="out/images",
    annotations_directory_path="out/annotations",
)

9.3 切分与合并

# 训练集/测试集切分(不修改原对象)
train_ds, test_ds = ds.split(split_ratio=0.8, random_state=42, shuffle=True)

# 再从 train 里切验证集
train_ds, valid_ds = train_ds.split(split_ratio=0.9, random_state=42)

# 合并多个数据集(自动统一类别表、重映射 class_id)
merged = sv.DetectionDataset.merge([ds_a, ds_b, ds_c])

merge() 会自动处理类别 ID 冲突:A 数据集的 cat=0、B 数据集的 dog=0, 合并后会重排成 cat=0, dog=1。手工合并数据集最容易错的就是这一步,它替你做了。

9.4 遍历检查标注

for image_path, image, annotations in ds:
    annotated = sv.BoxAnnotator().annotate(image.copy(), annotations)
    cv2.imwrite(f"check/{Path(image_path).name}", annotated)

标注质检的标准做法:批量渲染出来肉眼扫一遍,比看 txt 文件靠谱得多。


10. 第七课:模型评估(mAP / 混淆矩阵)

10.1 mAP

import supervision as sv
from supervision.metrics import MeanAveragePrecision, MetricTarget

map_metric = MeanAveragePrecision(
    metric_target=MetricTarget.BOXES,    # 也可 MASKS / ORIENTED_BOUNDING_BOXES
    class_agnostic=False,
)

for image_path, image, targets in ds:          # targets = 真值
    results = model(image)[0]
    predictions = sv.Detections.from_ultralytics(results)
    map_metric.update(predictions, targets)

map_result = map_metric.compute()
print(f"mAP@50    : {map_result.map50}")
print(f"mAP@75    : {map_result.map75}")
print(f"mAP@50-95 : {map_result.map50_95}")

update() 返回自身,所以单次也可以链式写:

map_result = MeanAveragePrecision().update(predictions, targets).compute()

0.26.1 起 MeanAveragePrecision 与 pycocotools 口径对齐, 数字可以直接和论文/榜单比较。

10.2 混淆矩阵

confusion_matrix = sv.ConfusionMatrix.from_detections(
    predictions=predictions_list,      # list[Detections]
    targets=targets_list,              # list[Detections]
    classes=ds.classes,
)
confusion_matrix.plot()                # 直接出图

这是找模型弱点最快的方法:一眼看出哪两类被混淆、哪类漏检最多。

10.3 其它可用指标

Precision、Recall、MeanAverageRecall、F1Score、IntersectionOverUnion (from supervision.metrics import ...),用法与 mAP 一致:update() → compute()。


11. 综合实战:路口车流统计

把前面所有零件拼起来 —— 一个可交付的完整程序:

"""路口车流统计:分方向计数 + 轨迹可视化"""
import numpy as np
import supervision as sv
from ultralytics import YOLO

SOURCE = "traffic.mp4"
TARGET = "traffic_out.mp4"
VEHICLE_CLASSES = [2, 3, 5, 7]      # COCO: car, motorcycle, bus, truck

model = YOLO("yolov8n.pt")
video_info = sv.VideoInfo.from_video_path(SOURCE)

# --- 组件全部在循环外创建 ---
tracker = sv.ByteTrack(frame_rate=video_info.fps)
smoother = sv.DetectionsSmoother()

line_zone = sv.LineZone(
    start=sv.Point(0, video_info.height // 2),
    end=sv.Point(video_info.width, video_info.height // 2),
)

box_annotator   = sv.BoxAnnotator(color_lookup=sv.ColorLookup.TRACK)
label_annotator = sv.LabelAnnotator(color_lookup=sv.ColorLookup.TRACK)
trace_annotator = sv.TraceAnnotator(
    color_lookup=sv.ColorLookup.TRACK, trace_length=50
)
line_annotator  = sv.LineZoneAnnotator(text_scale=1.0)


def callback(frame: np.ndarray, index: int) -> np.ndarray:
    results = model(frame, verbose=False)[0]
    detections = sv.Detections.from_ultralytics(results)

    # 过滤:只要车,且置信度够高
    detections = detections[
        np.isin(detections.class_id, VEHICLE_CLASSES)
        & (detections.confidence > 0.4)
    ]

    # 追踪 + 平滑
    detections = tracker.update_with_detections(detections)
    detections = smoother.update_with_detections(detections)

    # 计数
    line_zone.trigger(detections)

    # 可视化
    labels = [
        f"#{tid} {name}"
        for tid, name in zip(detections.tracker_id, detections.data["class_name"])
    ]
    annotated = trace_annotator.annotate(frame.copy(), detections=detections)
    annotated = box_annotator.annotate(annotated, detections=detections)
    annotated = label_annotator.annotate(annotated, detections=detections, labels=labels)
    annotated = line_annotator.annotate(annotated, line_counter=line_zone)
    return annotated


sv.process_video(source_path=SOURCE, target_path=TARGET, callback=callback)

print(f"南向北: {line_zone.in_count} 辆")
print(f"北向南: {line_zone.out_count} 辆")
print("分车型:", line_zone.in_count_per_class)

这段代码体现的工程范式(可以直接套到别的项目):

  1. 读视频元信息 —— 用真实分辨率 / 帧率初始化组件
  2. 所有有状态的组件在循环外创建(tracker / zone / annotator)—— 最容易错的一步
  3. callback 里的顺序固定:推理 → 过滤 → 追踪 → 平滑 → 判定 → 标注
  4. 标注顺序 = 图层顺序,先画的在底层(轨迹 → 框 → 文字 → 计数线)

12. 常见坑与注意事项

#坑后果正解
1ByteTrack() 建在 callback 里面追踪 ID 每帧重置,全部失效一定建在循环外
2LabelAnnotator 写中文显示 ???用 RichLabelAnnotator + 指定中文字体路径
3没传 frame.copy()原始帧被就地改写,多个 annotator 相互污染第一个 annotator 传 .copy()
4过滤条件用 andnumpy 报 truth value of array is ambiguous用 & 并给每个条件加括号
5LineZone 不追踪就用计数全是 0 或乱跳必须先 update_with_detections()
6PolygonZone 锚点选错目标在区外却被计数行人用 BOTTOM_CENTER
7InferenceSlicer 用于实时视频帧率掉到 1 fps 以下只用于离线大图;实时场景改用高分辨率输入或分区推理
8直接改 detections.xyxy 的切片有时是视图有时是拷贝,行为不一致用 detections[mask] 生成新对象
9版本升级不看 CHANGELOG0.x 版本 API 会变(如 BoxAnnotator 曾拆分过)生产环境锁定版本号 supervision==0.30.1
10以为它能训练模型找不到 train()它只做推理后处理,训练归 Ultralytics/Roboflow

第 9 条特别重要:supervision 仍是 0.x 版本,语义化版本约定下 次版本号变更允许破坏兼容。 项目里请写死版本,升级前先读 Releases。


13. 学习路径建议

三个阶段

Supervision 三阶段学习路径 按真实耗时排布的时间轴:半天跑通静态图片检测与标注,再用一天掌握视频处理与 ByteTrack 追踪,最后两天做越线计数、区域告警与模型评估等业务功能,累计约三天半。 0 1 2 3 4 天 阶段一 · 半天 静态图片检测 + 标注 验收 · 能自定义标签 阶段三 · 两天 计数 · 区域 · 评估 验收 · 完成车流统计 阶段二 · 一天 process_video + ByteTrack 验收 · 同一人全程一个 ID

官方资源

资源地址
文档主站https://supervision.roboflow.com/latest/
GitHub 仓库https://github.com/roboflow/supervision
Release 记录https://github.com/roboflow/supervision/releases
Cookbook(可运行 Notebook)https://supervision.roboflow.com/latest/cookbooks/
多边形绘制工具https://polygonzone.roboflow.com
配套项目 notebookshttps://github.com/roboflow/notebooks

生态位置

Roboflow 的开源矩阵,Supervision 是其中的「工具层」:

项目职责
supervision后处理工具箱(本文)
inference推理服务器,一行命令部署模型
rf-detrRoboflow 自研的 DETR 系检测模型
autodistill用大模型自动标注,蒸馏出小模型
maestro多模态大模型(VLM)微调
notebooks各类 CV 任务的教学 Notebook

系列文章


附:一页速查表

import supervision as sv

# ── 数据结构 ──────────────────────────────
det = sv.Detections.from_ultralytics(results)
det = det[det.confidence > 0.5]
det = det[det.class_id == 0]
det = det.with_nms(threshold=0.45)

# ── 标注 ─────────────────────────────────
sv.BoxAnnotator().annotate(scene=img.copy(), detections=det)
sv.LabelAnnotator().annotate(scene=img, detections=det, labels=labels)
sv.MaskAnnotator().annotate(scene=img, detections=det)
sv.TraceAnnotator().annotate(scene=img, detections=det)

# ── 视频 ─────────────────────────────────
info = sv.VideoInfo.from_video_path("in.mp4")
for frame in sv.get_video_frames_generator("in.mp4"): ...
sv.process_video("in.mp4", "out.mp4", callback=callback)

# ── 追踪 ─────────────────────────────────
tracker = sv.ByteTrack()               # 循环外!
det = tracker.update_with_detections(det)

# ── 计数 ─────────────────────────────────
line = sv.LineZone(sv.Point(0, 500), sv.Point(1920, 500))
line.trigger(det); line.in_count; line.out_count

zone = sv.PolygonZone(polygon=poly)
mask = zone.trigger(det); zone.current_count

# ── 大图切片 ─────────────────────────────
slicer = sv.InferenceSlicer(callback=cb, slice_wh=640, overlap_wh=100)
det = slicer(image)

# ── 数据集 ───────────────────────────────
ds = sv.DetectionDataset.from_coco(imgs, ann_json)
ds.as_yolo(out_imgs, out_labels, out_yaml)
train, test = ds.split(0.8)

# ── 评估 ─────────────────────────────────
from supervision.metrics import MeanAveragePrecision
r = MeanAveragePrecision().update(pred, gt).compute()
r.map50, r.map50_95