vision3d.ops#
Geometric operators for 3D data.
Functions
|
Accumulate and time-stamp a set of lidar sweeps. |
|
Class-aware 3D NMS: runs |
|
Convert 3D bounding boxes from |
|
Compute the 8 world-space corners of 3D bounding boxes. |
|
Compute the pairwise intersection-over-union of 3D |
|
Check 3D overlap between two sets of oriented bounding boxes. |
|
Decompose 3D boxes into centers, half-dimensions, and rotation matrices. |
|
Greedy, class-agnostic non-maximum suppression on 3D bounding boxes. |
|
Compute a boolean mask indicating which points fall inside which boxes. |
|
Return per-point box assignment. |
|
Return a mask selecting 3D points that project inside an image. |
|
Project 3D points in lidar frame to pixel coordinates. |
|
Bucket points into a 3D voxel grid. |
- vision3d.ops.accumulate_sweeps(sweeps, transforms, time_offsets)[source]#
Accumulate and time-stamp a set of lidar sweeps.
Each sweep is mapped into a common target frame by its own rigid transform, then all sweeps are concatenated into a single point cloud with a new trailing column holding the per-point time offset. This densifies a sparse lidar frame by folding in neighbouring sweeps while recording when each point was captured.
The transform is applied to the
(x, y, z)coordinates only. Feature columns (e.g. intensity) pass through unchanged and the time offset is appended after them, so a sweep of shape[N, 3+C]contributes rows of shape[N, 3+C+1].- Parameters:
sweeps (Sequence[Tensor]) – Sweeps to aggregate, each a
[N_i, 3+C]point cloud whose first three columns are(x, y, z)in that sweep’s own frame, followed byCfeature columns. Every sweep must share the same number of feature columns.transforms (Tensor) – Rigid
[S, 4, 4]homogeneous transforms, one per sweep, mapping that sweep’s coordinates into the target frame.time_offsets (Tensor) –
[S]per-sweep time offsets (e.g. seconds relative to the target frame) broadcast into the appended column.
- Returns:
Aggregated
[sum(N_i), 3+C+1]point cloud in the target frame, with the time offset as the last column. Rows follow the order ofsweeps.- Raises:
ValueError – If
sweepsis empty, or iftransformsortime_offsetsdo not have exactly one entry per sweep.- Return type:
- vision3d.ops.batched_nms_3d(boxes, scores, idxs, iou_threshold, format)[source]#
Class-aware 3D NMS: runs
nms_3d()independently per class.- Parameters:
- Returns:
int64tensor of indices intoboxesthat survived, sorted in decreasing order of score.- Return type:
- vision3d.ops.box3d_convert(boxes, in_fmt, out_fmt)[source]#
Convert 3D bounding boxes from
in_fmttoout_fmt.Only the lossless
XYZXYZ<->XYZLWHconversion is supported. All other conversions would discard or fabricate rotation angles.- Parameters:
boxes (Tensor) – Boxes to convert with shape
[..., K]. Supports any number of leading batch dimensions.in_fmt (BoundingBox3DFormat | str) – Source format.
out_fmt (BoundingBox3DFormat | str) – Target format.
- Returns:
Converted boxes with the same leading dimensions.
- Return type:
- vision3d.ops.box3d_corners(boxes, format)[source]#
Compute the 8 world-space corners of 3D bounding boxes.
Supports all rotation formats including full 9-DOF (yaw, pitch, roll).
Corner ordering:
4 -------- 5 z x |\ |\ | / | 7 -------| 6 | / | | | | |/ 0 |--------1 | +------ y \| \| 3 -------- 2
Bottom face (z-): {0, 1, 2, 3}. Top face (z+): {4, 5, 6, 7}.
- Parameters:
boxes (Tensor) – 3D bounding boxes
[N, K].format (BoundingBox3DFormat) – Format of the bounding boxes.
- Returns:
Corner coordinates
[N, 8, 3].- Return type:
- vision3d.ops.box3d_iou(boxes1, boxes2, format)[source]#
Compute the pairwise intersection-over-union of 3D
boxes1andboxes2.iou = vol / (vol1 + vol2 - vol), wherevolis the volume of the intersecting convex polyhedron andvol1,vol2are the volumes of the two input boxes.The same algorithm handles every supported box format — including full 9-DOF orientation (
XYZLWHYPR) — because the clipping step operates on the 8 box corners regardless of how they were produced.Note: This function is not differentiable.
- Parameters:
boxes1 (Tensor) – First set of boxes
[N, K].boxes2 (Tensor) – Second set of boxes
[M, K].format (BoundingBox3DFormat) – Format of both box sets.
- Returns:
[N, M]matrix of IoU values in[0, 1].- Return type:
- vision3d.ops.box3d_overlap(boxes1, boxes2, format)[source]#
Check 3D overlap between two sets of oriented bounding boxes.
Uses the Separating Axis Theorem (SAT) with 15 potential separating axes (3 face normals per box + 9 edge cross products).
- Parameters:
boxes1 (Tensor) – First set of boxes
[N, K].boxes2 (Tensor) – Second set of boxes
[M, K].format (BoundingBox3DFormat) – Format of both box sets.
- Returns:
Boolean matrix
[N, M]where True indicates overlap.- Return type:
- vision3d.ops.extract_box3d_params(boxes, format)[source]#
Decompose 3D boxes into centers, half-dimensions, and rotation matrices.
Supports all box formats including full 9-DOF (yaw, pitch, roll); formats without rotation yield identity rotation matrices.
- Parameters:
boxes (Tensor) – 3D bounding boxes
[M, K]whereKdepends onformat.format (BoundingBox3DFormat) – Format of the bounding boxes.
- Returns:
(centers, half_dims, rot)wherecentersandhalf_dimsare[M, 3]androtis[M, 3, 3]. Eachrotmaps local to world coordinates (world = rot @ local, matchingbox3d_corners), so a box’s world-frame axes are the columns ofrot. Transpose for the inverse (world-to-local) mapping.- Raises:
ValueError – If
formatis not a supported format.- Return type:
- vision3d.ops.nms_3d(boxes, scores, iou_threshold, format)[source]#
Greedy, class-agnostic non-maximum suppression on 3D bounding boxes.
Iteratively removes lower-scoring boxes whose IoU with a higher-scoring box exceeds
iou_threshold.- Parameters:
boxes (Tensor) –
[N, K]boxes to perform NMS on.Kdepends onformat.scores (Tensor) –
[N]prediction confidences.iou_threshold (float) – Discard any box whose IoU with a higher-scoring kept box is strictly greater than this value.
format (BoundingBox3DFormat) – Format of
boxes.
- Returns:
int64tensor of indices intoboxesthat survived, sorted in decreasing order of score.- Return type:
- vision3d.ops.points_in_boxes_3d(points, boxes, format)[source]#
Compute a boolean mask indicating which points fall inside which boxes.
Supports all rotation formats including full 9-DOF (yaw, pitch, roll).
- Parameters:
points (Tensor) – Point cloud coordinates
[N, 3+C]. Only the first 3 columns (x, y, z) are used.boxes (Tensor) – 3D bounding boxes
[M, K]where K depends on format.format (BoundingBox3DFormat) – Format of the bounding boxes.
- Returns:
Boolean tensor
[N, M]where entry(i, j)is True if pointiis inside boxj.- Return type:
- vision3d.ops.points_in_boxes_3d_indices(points, boxes, format)[source]#
Return per-point box assignment.
If a point is inside multiple boxes, the first (lowest index) box wins.
- Parameters:
points (Tensor) – Point cloud coordinates
[N, 3+C].boxes (Tensor) – 3D bounding boxes
[M, K].format (BoundingBox3DFormat) – Format of the bounding boxes.
- Returns:
Integer tensor
[N]with the index of the box each point belongs to, or-1if the point is not in any box.- Return type:
- vision3d.ops.points_in_image(points_3d, extrinsics, intrinsics, image_size)[source]#
Return a mask selecting 3D points that project inside an image.
- Parameters:
- Returns:
Boolean mask
[N]. An entry is true when the point has positive camera-frame depth and projects within the image bounds.- Return type:
- vision3d.ops.project_to_image(points_3d, extrinsics, intrinsics)[source]#
Project 3D points in lidar frame to pixel coordinates.
- Parameters:
- Returns:
(uv, depth)whereuvis the pixel coordinates[N, 2](u, v) anddepthis the camera-frame depth[N].- Return type:
- vision3d.ops.voxelize(points, point_cloud_range, voxel_size, max_points_per_voxel=32, max_voxels=None)[source]#
Bucket points into a 3D voxel grid.
For PointPillars-style pillars set
voxel_size = (dx, dy, z_max - z_min)so the grid has a single z-slice and every kept point lands atiz = 0. For 3D voxel detectors (VoxelNet, SECOND, CenterPoint-Voxel) use a normal three-axisvoxel_size.The op is not differentiable.
Grid dimensions are computed as
round((max - min) / voxel_size)with ties going away from zero. Ranges that are not an exact multiple ofvoxel_sizeround to the nearest integer cell count. Points landing in the partial trailing cell are clamped to the final grid index. Points with non-finite coordinates (NaN, +/-Inf) are treated as out-of-range and dropped.- Parameters:
points (Tensor) –
[N, C]float32 point cloud. The first three columns must be(x, y, z). Remaining columns may be arbitrary per-point features. Points outsidepoint_cloud_rangealong any axis are discarded.point_cloud_range (tuple[float, float, float, float, float, float]) –
(x_min, y_min, z_min, x_max, y_max, z_max).voxel_size (tuple[float, float, float]) –
(dx, dy, dz)voxel size.max_points_per_voxel (int) – Cap on points stored per voxel. Surplus points are dropped in input order. Default:
32.max_voxels (int | None) – Optional cap on the number of output voxels.
None(default) means no cap. When set, points that would have created a new voxel beyondmax_voxelsare silently dropped. Points landing in already-allocated voxels still get written.
- Returns:
(voxels, coords, num_points)wherevoxelsis[P, max_points_per_voxel, C]per-voxel point buffers padded with zeros at unfilled slots,coordsis[P, 3]int64(iz, iy, ix)voxel indices, andnum_pointsis[P]int64point counts per voxel, each in[1, max_points_per_voxel].Pis the number of non-empty voxels (capped atmax_voxelswhen set). When no points fall insidepoint_cloud_range, all three tensors have a leading dimension of zero.- Return type: