SAM3 fine-tune β cap_20260716_115414 (task8, segmentation)
Segmentation-head fine-tune of facebook/sam3,
trained end-to-end (box + mask losses) on 450 train / 50 val pseudo-RGB
frames from a static roadside Ouster OS1-128 capture (cap_20260716_115414,
6107 frames @ 10Hz), curated from CVAT task 8 annotations.
- Checkpoint: epoch 18 (of 20), selected by best validation bbox AP.
- Validation: AP=0.573, AP50=0.774, AP75=0.690.
- Classes (SAM3 text prompts):
person, bicycle, motorcycle, car, bus, truck. - Format: flat
SAM3Imagestate_dict (1132 keys, incl. backbone + trainedsegmentation_head.*), not nested under a"detector."prefix like Meta's official release checkpoints β load withcheckpoint_path=Nonethenmodel.load_state_dict(state_dict, strict=False)manually, not viabuild_sam3_image_model's owncheckpoint_path=loader.
Loss (Sam3LossWrapper, matcher = BinaryHungarianMatcherV2 + o2m
BinaryOneToManyMatcher): Boxes loss_bbox=5.0/loss_giou=2.0; IABCEMdetr
loss_ce=20.0/presence_loss=20.0/pos_weight=10/focal Ξ±=0.25,Ξ³=2; Masks
loss_mask=200.0/loss_dice=10.0, point-sampled (12544 points,
oversample_ratio=3.0, importance_sample_ratio=0.75), final decoder stage only.
Part of a larger pipeline (2D SAM3 segmentation β 3D box reconstruction β tracking-based refinement β SUSTechPOINTS export) for this LiDAR capture.
Model tree for zuhdifr/sam3-cheonan-sicheong
Base model
facebook/sam3