Supervision:把模型输出变成能交付的业务功能
Roboflow 出品的计算机视觉「胶水层」。它不训练模型也不做推理,专门把各家模型吐出来的原始结果,变成越线计数、区域告警、轨迹追踪这些真能交付的东西
做视觉项目最烦的一件事:模型换一次,后处理代码就得重写一遍。YOLO 返回一种结构,RF-DETR 返回另一种,Transformers 又是第三种。
Supervision 干的就是这件事——把二十多家推理框架的输出洗成同一个数据结构,然后在上面配齐一整套业务工具:画框、追踪 ID、越线计数、区域人数、数据集格式互转、mAP 评估。
| 版本 | GitHub Star | PyPI 月下载 | 协议 | Python |
|---|---|---|---|---|
| 0.30.1 | 38,000+ | 100 万+ | MIT(商用无限制) | ≥ 3.10 |
下面是一份可以照着跑的完整教程,从装环境到「路口车流统计」这种可交付的成品。
1. 它到底是什么
一句话:Supervision 不训练模型、不做推理,它把「模型吐出来的原始结果」变成「能用的业务功能」。
它在整个流程中的位置
为什么需要它
没有 Supervision 时,你换一次模型(YOLO → RF-DETR)就得重写一遍后处理代码, 因为每家返回的结果结构都不一样。有了它:
detections = sv.Detections.from_ultralytics(results) # 换模型
detections = sv.Detections.from_transformers(results) # 只改这一行
# ↓ 后面所有代码完全不变
这就是它最大的价值:解耦。 模型可以随便换,业务代码一行不动。
项目现状(2026-09)
| 项 | 值 |
|---|---|
| 最新版本 | 0.30.1 |
| GitHub Star | 38,000+ |
| PyPI 月下载量 | 100 万+ |
| 协议 | MIT(商用无限制) |
| Python | ≥ 3.10 |
| 维护方 | Roboflow(商业公司,持续维护) |
2. 安装
pip install supervision
配套装一个推理框架(教程用 Ultralytics YOLO 举例):
pip install ultralytics
验证:
import supervision as sv
print(sv.__version__)
约定:社区习惯
import supervision as sv,本文全部沿用。
3. 核心概念:Detections
整个库只有一个核心数据结构 —— sv.Detections。理解它 = 掌握 80% 的库。
3.1 它长什么样
sv.Detections(
xyxy, # (n, 4) float 必填 每个框的 [x1, y1, x2, y2]
mask, # (n, H, W) bool 可选 分割掩码
confidence, # (n,) float 可选 置信度
class_id, # (n,) int 可选 类别编号
tracker_id, # (n,) int 可选 追踪 ID(追踪后才有)
data, # dict 可选 自定义字段,如 class_name
metadata, # dict 可选 集合级信息,如视频名、时间戳
)
关键理解:它本质是一堆 numpy 数组的容器,n 是检测到的目标个数。 所以它天然支持 numpy 式的过滤和切片。
3.2 从各家模型导入(from_* 方法)
| 推理来源 | 方法 |
|---|---|
| Ultralytics (YOLOv8/v11) | sv.Detections.from_ultralytics(results) |
| Roboflow Inference | sv.Detections.from_inference(results) |
| HuggingFace Transformers | sv.Detections.from_transformers(results) |
| RF-DETR | 直接返回 Detections,无需转换 |
| YOLOv5 | sv.Detections.from_yolov5(results) |
| YOLO-NAS | sv.Detections.from_yolo_nas(results) |
| Detectron2 | sv.Detections.from_detectron2(results) |
| MMDetection | sv.Detections.from_mmdetection(results) |
| SAM / SAM3 | sv.Detections.from_sam() / from_sam3() |
| PaddleDetection | sv.Detections.from_paddledet(results) |
| TensorFlow | sv.Detections.from_tensorflow(results) |
| DeepSparse / NCNN | from_deepsparse() / from_ncnn() |
| EasyOCR | sv.Detections.from_easyocr(results) |
| Azure AI Vision | sv.Detections.from_azure_analyze_image(results) |
| 大模型 VLM/LMM | sv.Detections.from_vlm() / from_lmm() |
3.3 过滤:像操作 DataFrame 一样
# 只保留置信度 > 0.5 的
detections = detections[detections.confidence > 0.5]
# 只保留 class_id == 0(比如"人")
detections = detections[detections.class_id == 0]
# 只保留框面积 > 1000 像素的
detections = detections[detections.box_area > 1000]
# 多条件组合(注意用 & 不是 and,用括号包起来)
detections = detections[
(detections.confidence > 0.5) & (detections.class_id == 0)
]
# 按数量
print(len(detections))
3.4 常用方法
detections.with_nms(threshold=0.45) # 非极大值抑制,去重叠框
detections.with_soft_nms(sigma=0.5) # Soft-NMS
detections.with_nmm(threshold=0.5) # 非极大值合并(融合而非丢弃)
detections.area # 掩码面积(有 mask 时)
detections.box_area # 框面积
detections.box_aspect_ratio # 宽高比
detections.get_anchors_coordinates(sv.Position.BOTTOM_CENTER) # 取锚点坐标
detections.get_data("class_name") # 取自定义字段
sv.Detections.merge([det_a, det_b]) # 合并多组检测
sv.Detections.empty() # 空检测(占位用)
detections.is_empty() # 判空
get_anchors_coordinates很重要:越线/区域判定都是靠「用框上的哪个点代表这个物体」。 行人通常用BOTTOM_CENTER(脚底),车辆有时用CENTER。
4. 第一课:图片检测 + 标注
4.1 最小可运行例子
import cv2
import supervision as sv
from ultralytics import YOLO
# 1) 跑模型
model = YOLO("yolov8n.pt")
image = cv2.imread("input.jpg")
results = model(image)[0]
# 2) 转成 Supervision 的统一结构
detections = sv.Detections.from_ultralytics(results)
# 3) 标注
box_annotator = sv.BoxAnnotator()
label_annotator = sv.LabelAnnotator()
annotated = box_annotator.annotate(scene=image.copy(), detections=detections)
annotated = label_annotator.annotate(scene=annotated, detections=detections)
cv2.imwrite("output.jpg", annotated)
注意三个模式:
- Annotator 是先创建对象、再反复调用
.annotate()(性能考虑,别在循环里 new) .annotate()返回图像,所以可以链式叠加- 传
image.copy()而不是image,避免原图被就地改写
4.2 自定义标签
默认标签是「类别名 置信度」。想自定义:
labels = [
f"{class_name} {confidence:.2f}"
for class_name, confidence
in zip(detections["class_name"], detections.confidence)
]
annotated = label_annotator.annotate(
scene=annotated, detections=detections, labels=labels
)
4.3 Annotator 全家桶(30+ 个)
按用途分五类,都可以叠加使用:
| 类别 | Annotator | 效果 |
|---|---|---|
| 轮廓类 | BoxAnnotator | 标准矩形框 |
RoundBoxAnnotator | 圆角框 | |
BoxCornerAnnotator | 只画四个角(科技感) | |
OrientedBoxAnnotator | 旋转框(OBB,如遥感、文字) | |
CircleAnnotator / EllipseAnnotator | 圆 / 椭圆(体育赛事常用) | |
PolygonAnnotator | 多边形轮廓 | |
| 填充类 | ColorAnnotator | 半透明色块填充 |
MaskAnnotator | 实例分割掩码(需要 detections.mask) | |
HaloAnnotator | 光晕效果 | |
| 标记类 | DotAnnotator | 中心点圆点 |
TriangleAnnotator | 三角标记(指向目标) | |
| 文字类 | LabelAnnotator | 带背景的文字标签 |
RichLabelAnnotator | 自定义字体(中文标签必须用这个) | |
| 图像处理类 | BlurAnnotator | 模糊目标区域(打码,如人脸/车牌) |
PixelateAnnotator | 马赛克 | |
| 追踪/聚合类 | TraceAnnotator | 运动轨迹线(需要 tracker_id) |
HeatMapAnnotator | 热力图(人流密度分析) | |
PercentageBarAnnotator | 进度条形式显示置信度 | |
| 其它 | IconAnnotator | 贴自定义图标 |
BackgroundOverlayAnnotator | 背景变暗,突出目标 | |
ComparisonAnnotator | 两组检测结果对比(调参神器) | |
CropAnnotator | 把目标裁剪放大显示在旁边 |
中文标签的坑:
LabelAnnotator底层用 OpenCV 画字,不支持中文,会显示成???。 中文必须用RichLabelAnnotator并指定字体文件:label_annotator = sv.RichLabelAnnotator(font_path="C:/Windows/Fonts/msyh.ttc")
4.4 颜色控制
box_annotator = sv.BoxAnnotator(
color=sv.ColorPalette.DEFAULT,
color_lookup=sv.ColorLookup.CLASS, # 按类别上色(默认)
# sv.ColorLookup.TRACK → 按追踪 ID 上色(每个目标一个固定颜色)
# sv.ColorLookup.INDEX → 按序号上色
thickness=2,
)
做追踪演示时,把
color_lookup设成TRACK,同一个人全程一个颜色,观感立刻专业。
5. 第二课:视频处理与目标追踪
5.1 视频工具三件套
import supervision as sv
# 读元信息
video_info = sv.VideoInfo.from_video_path("input.mp4")
print(video_info.width, video_info.height, video_info.fps, video_info.total_frames)
# 逐帧生成器(省内存,不会一次读进来)
for frame in sv.get_video_frames_generator("input.mp4"):
...
# 只取部分帧
for frame in sv.get_video_frames_generator("input.mp4", start=100, end=300, stride=2):
...
5.2 sv.process_video:视频处理的标准范式
这是全库最实用的一个函数,帮你搞定「读帧 → 处理 → 写回」的全部样板代码:
import numpy as np
import supervision as sv
from ultralytics import YOLO
model = YOLO("yolov8n.pt")
box_annotator = sv.BoxAnnotator()
def callback(frame: np.ndarray, index: int) -> np.ndarray:
results = model(frame)[0]
detections = sv.Detections.from_ultralytics(results)
return box_annotator.annotate(frame.copy(), detections=detections)
sv.process_video(
source_path="input.mp4",
target_path="output.mp4",
callback=callback
)
范式要点:你只需要写 callback(帧, 帧号) -> 处理后的帧,其余全部交给它。
分辨率、帧率、编码器都自动继承源视频。
5.3 目标追踪:ByteTrack
检测只知道「这一帧有 3 个人」;追踪才知道「这 3 个人跟上一帧是不是同一批」。 Supervision 内置了 ByteTrack(不需要额外装库):
import numpy as np
import supervision as sv
from ultralytics import YOLO
model = YOLO("yolov8n.pt")
tracker = sv.ByteTrack() # ← 追踪器,全程只创建一次
box_annotator = sv.BoxAnnotator()
label_annotator = sv.LabelAnnotator()
def callback(frame: np.ndarray, _: int) -> np.ndarray:
results = model(frame)[0]
detections = sv.Detections.from_ultralytics(results)
detections = tracker.update_with_detections(detections) # ← 关键这一行
labels = [
f"#{tracker_id} {class_name}"
for class_name, tracker_id
in zip(detections.data["class_name"], detections.tracker_id)
]
annotated = box_annotator.annotate(frame.copy(), detections=detections)
return label_annotator.annotate(annotated, detections=detections, labels=labels)
sv.process_video("people-walking.mp4", "result.mp4", callback=callback)
调用 update_with_detections() 后,detections.tracker_id 就被填上了。
必须注意:
tracker = sv.ByteTrack()一定要放在 callback 外面。 放里面等于每帧重建追踪器,ID 会一直从 1 开始跳,追踪完全失效 —— 这是最高频的新手错误。
5.4 加轨迹线
trace_annotator = sv.TraceAnnotator(trace_length=30) # 保留最近 30 帧的轨迹
def callback(frame, _):
results = model(frame)[0]
detections = sv.Detections.from_ultralytics(results)
detections = tracker.update_with_detections(detections)
annotated = box_annotator.annotate(frame.copy(), detections=detections)
return trace_annotator.annotate(annotated, detections=detections)
5.5 抖动平滑
框跳动厉害时,加一个平滑器:
smoother = sv.DetectionsSmoother()
detections = tracker.update_with_detections(detections)
detections = smoother.update_with_detections(detections) # 追踪之后
注意顺序:先追踪、后平滑。平滑器依赖 tracker_id 来判断哪些框属于同一个目标。
6. 第三课:越线计数(LineZone)
典型场景:门口进出人数、路口车流量、传送带计件。
6.1 原理
画一条虚拟线,物体的锚点从线的一侧穿到另一侧就计一次数。 必须先有追踪 ID(不然没法判断”同一个物体”跨越了)。
6.2 API
sv.LineZone(
start: sv.Point,
end: sv.Point,
triggering_anchors = ( # 用框上哪些点判定穿越
sv.Position.TOP_LEFT,
sv.Position.TOP_RIGHT,
sv.Position.BOTTOM_LEFT,
sv.Position.BOTTOM_RIGHT,
),
minimum_crossing_threshold: int = 1, # 需连续多少帧确认,抗抖动
)
关键属性:
| 属性 | 含义 |
|---|---|
in_count | 从外向内穿越的累计数 |
out_count | 从内向外穿越的累计数 |
in_count_per_class | 分类别的进入计数(dict) |
out_count_per_class | 分类别的离开计数(dict) |
6.3 完整示例
import numpy as np
import supervision as sv
from ultralytics import YOLO
model = YOLO("yolov8n.pt")
tracker = sv.ByteTrack()
# 定义计数线(坐标是像素,按你的视频分辨率来定)
line_zone = sv.LineZone(
start=sv.Point(x=0, y=500),
end=sv.Point(x=1920, y=500),
)
line_annotator = sv.LineZoneAnnotator(
thickness=2,
text_thickness=2,
text_scale=1.0,
custom_in_text="进", # 中文这里同样受 OpenCV 限制,英文更稳
custom_out_text="出",
)
box_annotator = sv.BoxAnnotator()
def callback(frame: np.ndarray, _: int) -> np.ndarray:
results = model(frame)[0]
detections = sv.Detections.from_ultralytics(results)
detections = detections[detections.class_id == 0] # 只算"人"
detections = tracker.update_with_detections(detections)
line_zone.trigger(detections) # ← 更新计数
annotated = box_annotator.annotate(frame.copy(), detections=detections)
return line_annotator.annotate(annotated, line_counter=line_zone)
sv.process_video("input.mp4", "output.mp4", callback=callback)
print(f"进入: {line_zone.in_count}, 离开: {line_zone.out_count}")
trigger()返回(crossed_in, crossed_out)两个布尔数组, 需要”谁穿过去了”的明细时接住它;只要总数的话直接调用不接返回值即可。
7. 第四课:区域计数(PolygonZone)
典型场景:某区域实时人数、车位占用、危险区域闯入告警。
7.1 API
sv.PolygonZone(
polygon: np.ndarray, # (N, 2) 的顶点坐标
triggering_anchors = (sv.Position.BOTTOM_CENTER,), # 默认脚底中心
require_all_anchors: bool = True, # True=所有锚点都在内才算;False=任一即可
)
| 成员 | 说明 |
|---|---|
trigger(detections) | 返回布尔数组,标记哪些检测在区域内 |
current_count | 当前区域内目标数 |
mask | 区域的二维布尔掩码 |
7.2 完整示例
import numpy as np
import supervision as sv
from ultralytics import YOLO
model = YOLO("yolov8n.pt")
polygon = np.array([[300, 200], [900, 200], [900, 800], [300, 800]])
zone = sv.PolygonZone(polygon=polygon)
zone_annotator = sv.PolygonZoneAnnotator(
zone=zone,
color=sv.Color.RED,
thickness=4,
text_scale=1.5,
opacity=0.2, # 区域半透明填充
)
box_annotator = sv.BoxAnnotator()
def callback(frame: np.ndarray, _: int) -> np.ndarray:
results = model(frame)[0]
detections = sv.Detections.from_ultralytics(results)
mask = zone.trigger(detections=detections) # ← 布尔数组
detections = detections[mask] # 只保留区域内的
annotated = box_annotator.annotate(frame.copy(), detections=detections)
annotated = zone_annotator.annotate(scene=annotated)
if zone.current_count > 10:
print("告警:区域人数超限")
return annotated
sv.process_video("input.mp4", "output.mp4", callback=callback)
triggering_anchors怎么选:行人用BOTTOM_CENTER(脚踩在哪算在哪), 空中/俯视目标用CENTER。选错会导致”人明明在区域外但框的角伸进来了也被计数”。
7.3 多边形怎么画出来
手写坐标很痛苦,用 Roboflow 的在线工具:https://polygonzone.roboflow.com 上传一张截图,鼠标点几下,直接生成 numpy 数组代码。
8. 第五课:小目标检测(InferenceSlicer)
8.1 问题
4K 航拍图直接丢进 YOLO,模型内部会把图缩到 640×640, 原本 20 像素的小车缩成 3 像素 —— 检不出来。
8.2 方案:切片推理(SAHI)
把大图切成若干 640×640 的小块,分别推理,再把结果拼回原图坐标系并去重。
8.3 完整签名
sv.InferenceSlicer(
callback, # 必填:单块推理函数
slice_wh = 640, # 切片尺寸
overlap_wh = 100, # 切片重叠像素
overlap_filter = OverlapFilter.NON_MAX_SUPPRESSION, # 拼接去重策略
iou_threshold = 0.5,
overlap_metric = OverlapMetric.IOU,
thread_workers = 1, # 并行线程数
compact_masks = False,
batch_size = 1,
)
8.4 用法
import cv2
import supervision as sv
from ultralytics import YOLO
model = YOLO("yolo11m.pt")
def callback(tile) -> sv.Detections:
results = model(tile)[0]
return sv.Detections.from_ultralytics(results)
slicer = sv.InferenceSlicer(
callback=callback,
slice_wh=640,
overlap_wh=100,
thread_workers=4, # 有 GPU 显存余量时开并行提速
)
image = cv2.imread("aerial.jpg")
detections = slicer(image) # 像调函数一样调用
调参经验:
- 目标经常在切片边界被切断 → 加大
overlap_wh - 太慢 → 减小
overlap_wh或加大slice_wh - 代价:切片数量 = 面积/切片面积,4K 图约 30 块,推理耗时约 30 倍,不适合实时视频
9. 第六课:数据集管理与格式转换
sv.DetectionsDataset 把 YOLO / COCO / Pascal VOC 三种格式打通,
互转只要两行。
9.1 加载
import supervision as sv
# YOLO 格式
ds = sv.DetectionDataset.from_yolo(
images_directory_path="dataset/images",
annotations_directory_path="dataset/labels",
data_yaml_path="dataset/data.yaml",
is_obb=False, # 旋转框数据集设 True
show_progress=True,
)
# COCO 格式
ds = sv.DetectionDataset.from_coco(
images_directory_path="dataset/images",
annotations_path="dataset/_annotations.coco.json",
)
# Pascal VOC 格式
ds = sv.DetectionDataset.from_pascal_voc(
images_directory_path="dataset/images",
annotations_directory_path="dataset/annotations",
)
print(ds.classes) # 类别列表
print(len(ds)) # 图片数
9.2 导出(= 格式转换)
# COCO → YOLO,就这两步
ds = sv.DetectionDataset.from_coco(...)
ds.as_yolo(
images_directory_path="out/images",
annotations_directory_path="out/labels",
data_yaml_path="out/data.yaml",
)
# YOLO → COCO
ds.as_coco(
images_directory_path="out/images",
annotations_path="out/annotations.json",
)
# → Pascal VOC
ds.as_pascal_voc(
images_directory_path="out/images",
annotations_directory_path="out/annotations",
)
9.3 切分与合并
# 训练集/测试集切分(不修改原对象)
train_ds, test_ds = ds.split(split_ratio=0.8, random_state=42, shuffle=True)
# 再从 train 里切验证集
train_ds, valid_ds = train_ds.split(split_ratio=0.9, random_state=42)
# 合并多个数据集(自动统一类别表、重映射 class_id)
merged = sv.DetectionDataset.merge([ds_a, ds_b, ds_c])
merge()会自动处理类别 ID 冲突:A 数据集的cat=0、B 数据集的dog=0, 合并后会重排成cat=0, dog=1。手工合并数据集最容易错的就是这一步,它替你做了。
9.4 遍历检查标注
for image_path, image, annotations in ds:
annotated = sv.BoxAnnotator().annotate(image.copy(), annotations)
cv2.imwrite(f"check/{Path(image_path).name}", annotated)
标注质检的标准做法:批量渲染出来肉眼扫一遍,比看 txt 文件靠谱得多。
10. 第七课:模型评估(mAP / 混淆矩阵)
10.1 mAP
import supervision as sv
from supervision.metrics import MeanAveragePrecision, MetricTarget
map_metric = MeanAveragePrecision(
metric_target=MetricTarget.BOXES, # 也可 MASKS / ORIENTED_BOUNDING_BOXES
class_agnostic=False,
)
for image_path, image, targets in ds: # targets = 真值
results = model(image)[0]
predictions = sv.Detections.from_ultralytics(results)
map_metric.update(predictions, targets)
map_result = map_metric.compute()
print(f"mAP@50 : {map_result.map50}")
print(f"mAP@75 : {map_result.map75}")
print(f"mAP@50-95 : {map_result.map50_95}")
update() 返回自身,所以单次也可以链式写:
map_result = MeanAveragePrecision().update(predictions, targets).compute()
0.26.1 起
MeanAveragePrecision与pycocotools口径对齐, 数字可以直接和论文/榜单比较。
10.2 混淆矩阵
confusion_matrix = sv.ConfusionMatrix.from_detections(
predictions=predictions_list, # list[Detections]
targets=targets_list, # list[Detections]
classes=ds.classes,
)
confusion_matrix.plot() # 直接出图
这是找模型弱点最快的方法:一眼看出哪两类被混淆、哪类漏检最多。
10.3 其它可用指标
Precision、Recall、MeanAverageRecall、F1Score、IntersectionOverUnion
(from supervision.metrics import ...),用法与 mAP 一致:update() → compute()。
11. 综合实战:路口车流统计
把前面所有零件拼起来 —— 一个可交付的完整程序:
"""路口车流统计:分方向计数 + 轨迹可视化"""
import numpy as np
import supervision as sv
from ultralytics import YOLO
SOURCE = "traffic.mp4"
TARGET = "traffic_out.mp4"
VEHICLE_CLASSES = [2, 3, 5, 7] # COCO: car, motorcycle, bus, truck
model = YOLO("yolov8n.pt")
video_info = sv.VideoInfo.from_video_path(SOURCE)
# --- 组件全部在循环外创建 ---
tracker = sv.ByteTrack(frame_rate=video_info.fps)
smoother = sv.DetectionsSmoother()
line_zone = sv.LineZone(
start=sv.Point(0, video_info.height // 2),
end=sv.Point(video_info.width, video_info.height // 2),
)
box_annotator = sv.BoxAnnotator(color_lookup=sv.ColorLookup.TRACK)
label_annotator = sv.LabelAnnotator(color_lookup=sv.ColorLookup.TRACK)
trace_annotator = sv.TraceAnnotator(
color_lookup=sv.ColorLookup.TRACK, trace_length=50
)
line_annotator = sv.LineZoneAnnotator(text_scale=1.0)
def callback(frame: np.ndarray, index: int) -> np.ndarray:
results = model(frame, verbose=False)[0]
detections = sv.Detections.from_ultralytics(results)
# 过滤:只要车,且置信度够高
detections = detections[
np.isin(detections.class_id, VEHICLE_CLASSES)
& (detections.confidence > 0.4)
]
# 追踪 + 平滑
detections = tracker.update_with_detections(detections)
detections = smoother.update_with_detections(detections)
# 计数
line_zone.trigger(detections)
# 可视化
labels = [
f"#{tid} {name}"
for tid, name in zip(detections.tracker_id, detections.data["class_name"])
]
annotated = trace_annotator.annotate(frame.copy(), detections=detections)
annotated = box_annotator.annotate(annotated, detections=detections)
annotated = label_annotator.annotate(annotated, detections=detections, labels=labels)
annotated = line_annotator.annotate(annotated, line_counter=line_zone)
return annotated
sv.process_video(source_path=SOURCE, target_path=TARGET, callback=callback)
print(f"南向北: {line_zone.in_count} 辆")
print(f"北向南: {line_zone.out_count} 辆")
print("分车型:", line_zone.in_count_per_class)
这段代码体现的工程范式(可以直接套到别的项目):
- 读视频元信息 —— 用真实分辨率 / 帧率初始化组件
- 所有有状态的组件在循环外创建(tracker / zone / annotator)—— 最容易错的一步
- callback 里的顺序固定:推理 → 过滤 → 追踪 → 平滑 → 判定 → 标注
- 标注顺序 = 图层顺序,先画的在底层(轨迹 → 框 → 文字 → 计数线)
12. 常见坑与注意事项
| # | 坑 | 后果 | 正解 |
|---|---|---|---|
| 1 | ByteTrack() 建在 callback 里面 | 追踪 ID 每帧重置,全部失效 | 一定建在循环外 |
| 2 | LabelAnnotator 写中文 | 显示 ??? | 用 RichLabelAnnotator + 指定中文字体路径 |
| 3 | 没传 frame.copy() | 原始帧被就地改写,多个 annotator 相互污染 | 第一个 annotator 传 .copy() |
| 4 | 过滤条件用 and | numpy 报 truth value of array is ambiguous | 用 & 并给每个条件加括号 |
| 5 | LineZone 不追踪就用 | 计数全是 0 或乱跳 | 必须先 update_with_detections() |
| 6 | PolygonZone 锚点选错 | 目标在区外却被计数 | 行人用 BOTTOM_CENTER |
| 7 | InferenceSlicer 用于实时视频 | 帧率掉到 1 fps 以下 | 只用于离线大图;实时场景改用高分辨率输入或分区推理 |
| 8 | 直接改 detections.xyxy 的切片 | 有时是视图有时是拷贝,行为不一致 | 用 detections[mask] 生成新对象 |
| 9 | 版本升级不看 CHANGELOG | 0.x 版本 API 会变(如 BoxAnnotator 曾拆分过) | 生产环境锁定版本号 supervision==0.30.1 |
| 10 | 以为它能训练模型 | 找不到 train() | 它只做推理后处理,训练归 Ultralytics/Roboflow |
第 9 条特别重要:supervision 仍是 0.x 版本,语义化版本约定下 次版本号变更允许破坏兼容。 项目里请写死版本,升级前先读 Releases。
13. 学习路径建议
三个阶段
官方资源
| 资源 | 地址 |
|---|---|
| 文档主站 | https://supervision.roboflow.com/latest/ |
| GitHub 仓库 | https://github.com/roboflow/supervision |
| Release 记录 | https://github.com/roboflow/supervision/releases |
| Cookbook(可运行 Notebook) | https://supervision.roboflow.com/latest/cookbooks/ |
| 多边形绘制工具 | https://polygonzone.roboflow.com |
| 配套项目 notebooks | https://github.com/roboflow/notebooks |
生态位置
Roboflow 的开源矩阵,Supervision 是其中的「工具层」:
| 项目 | 职责 |
|---|---|
| supervision | 后处理工具箱(本文) |
| inference | 推理服务器,一行命令部署模型 |
| rf-detr | Roboflow 自研的 DETR 系检测模型 |
| autodistill | 用大模型自动标注,蒸馏出小模型 |
| maestro | 多模态大模型(VLM)微调 |
| notebooks | 各类 CV 任务的教学 Notebook |
系列文章
- 本篇 — Supervision 是什么、怎么用(教学文档)
- 换模型对照手册 — 16 家模型分别怎么接,附一览速查表
- 入门答疑:这些框架该选哪个 — 选型决策卡,以及 YOLO 许可证的坑
附:一页速查表
import supervision as sv
# ── 数据结构 ──────────────────────────────
det = sv.Detections.from_ultralytics(results)
det = det[det.confidence > 0.5]
det = det[det.class_id == 0]
det = det.with_nms(threshold=0.45)
# ── 标注 ─────────────────────────────────
sv.BoxAnnotator().annotate(scene=img.copy(), detections=det)
sv.LabelAnnotator().annotate(scene=img, detections=det, labels=labels)
sv.MaskAnnotator().annotate(scene=img, detections=det)
sv.TraceAnnotator().annotate(scene=img, detections=det)
# ── 视频 ─────────────────────────────────
info = sv.VideoInfo.from_video_path("in.mp4")
for frame in sv.get_video_frames_generator("in.mp4"): ...
sv.process_video("in.mp4", "out.mp4", callback=callback)
# ── 追踪 ─────────────────────────────────
tracker = sv.ByteTrack() # 循环外!
det = tracker.update_with_detections(det)
# ── 计数 ─────────────────────────────────
line = sv.LineZone(sv.Point(0, 500), sv.Point(1920, 500))
line.trigger(det); line.in_count; line.out_count
zone = sv.PolygonZone(polygon=poly)
mask = zone.trigger(det); zone.current_count
# ── 大图切片 ─────────────────────────────
slicer = sv.InferenceSlicer(callback=cb, slice_wh=640, overlap_wh=100)
det = slicer(image)
# ── 数据集 ───────────────────────────────
ds = sv.DetectionDataset.from_coco(imgs, ann_json)
ds.as_yolo(out_imgs, out_labels, out_yaml)
train, test = ds.split(0.8)
# ── 评估 ─────────────────────────────────
from supervision.metrics import MeanAveragePrecision
r = MeanAveragePrecision().update(pred, gt).compute()
r.map50, r.map50_95