Point Cloud Library (PCL) oneAPI Component#
PCL 1.15 oneAPI 2026.1+ SYCL 2020 Level Zero
The Point Cloud Library (PCL) oneAPI Optimized Component (libpcl-oneapi) delivers native Intel® oneAPI and SYCL™ 2020 acceleration for 3D point cloud and LiDAR processing pipelines in Intel® Robotics AI Suite for autonomous mobile robots (AMRs), automated guided vehicles (AGVs), and industrial vision systems.
It provides hardware acceleration across Intel® Core™ Ultra integrated GPUs, Intel® Arc™ discrete GPUs, and Intel® Xeon® processors, significantly reducing execution latency for perception tasks such as filtering, spatial indexing, feature extraction, surface reconstruction, and object segmentation.
1. Architectural Overview#
libpcl-oneapi is engineered as a modular acceleration layer that pairs with standard PCL architectures and the Intel® Robotics AI Suite ecosystem.
flowchart TD
subgraph AppLayer [Consuming Robotics AI Suite Application]
ROS2[ROS 2 Nodes / Perception Pipeline]
Nav2[Navigation2 / Costmaps / Localization]
end
subgraph PCLLayer [PCL Interface Layer]
PCLCore[Upstream PCL 1.15 API]
PCLOneAPI[libpcl-oneapi Acceleration Extension]
end
subgraph Accelerators [oneAPI Accelerated Subsystems]
Filters[pcl_oneapi_filters]
KdTree[pcl_oneapi_kdtree / search]
Features[pcl_oneapi_features]
SampleConsensus[pcl_oneapi_sample_consensus]
Segmentation[pcl_oneapi_segmentation]
Surface[pcl_oneapi_surface]
FLANN[libflann-oneapi Engine]
end
subgraph RuntimeLayer [Intel oneAPI Runtime]
DPC[DPC++ / SYCL 2020 Runtime]
L0[Level Zero Driver]
DPL[oneDPL / TBB]
end
subgraph HWLayer [Target Compute Devices]
iGPU[Intel Xe / Arc Integrated GPU]
dGPU[Intel Arc / Flex Discrete GPU]
CPU[Intel Xeon / Core CPU Fallback]
end
AppLayer --> PCLLayer
PCLOneAPI --> Accelerators
KdTree --> FLANN
Accelerators --> RuntimeLayer
RuntimeLayer --> HWLayer
Core Design Principles#
Built cleanly on modern SYCL 2020 standards and oneDPL parallel algorithms. All kernels use standardized sycl::queue, sycl::nd_item, sycl::sub_group, and oneapi::dpl execution policies for portable, high-performance hardware execution across Intel GPUs and CPUs.
Standalone Extension Mode (
PCL_ONEAPI_EXTENSION_ONLY=ON): Installs exclusively the accelerated extension libraries and headers against an existing distribution or ROS PCL installation (libpcl-dev).Full Integrated PCL Build (
BUILD_ONEAPI=ON): Provides a complete, unified PCL 1.15 distribution compiled end-to-end with Intel oneAPI compilers.
All accelerated modules expose .setQueue(sycl::queue&) methods accepting caller-owned sycl::queue instances (or automatically fall back to pcl::oneapi::getDefaultQueue()). Enables targeting specific GPU sub-devices, in-order queues, and overlapping asynchronous compute with sensor acquisition.
Supports persistent USM buffer reuse across frames. Continuous perception loops avoid host-to-device allocation thrashing by reusing pre-allocated device memory for repeated point cloud traversals.
libflann-oneapiHigh-dimensional and 3D spatial queries in pcl_oneapi_kdtree and pcl_oneapi_search delegate to libflann-oneapi (v1.9.2-oneAPI), delivering GPU-resident nearest neighbor searches.
2. Accelerated Subsystems#
libpcl-oneapi organizes hardware acceleration into modular shared libraries:
Library Target |
Subsystem |
Description & Acceleration Capabilities |
|---|---|---|
|
Common / Runtime |
Device selection, queue management ( |
|
Spatial Indexing |
3D KD-Tree search accelerated via Level Zero and |
|
Search Dispatcher |
Unified search interface for nearest-neighbor and radius queries. |
|
Octree Search |
Hierarchical octree construction, spatial voxel search, and change detection. |
|
Feature Estimation |
Surface normal estimation with parallel sub-group warp reductions. |
|
Filtering |
PassThrough, VoxelGrid, and StatisticalOutlierRemoval (SOR) downsampling. |
|
Alignment & ICP |
Iterative Closest Point (ICP), SAC-IA, and correspondence estimation. |
|
Model Fitting |
Parallel RANSAC and M-estimator plane, line, and cylinder model fitting. |
|
Segmentation |
Parallel sample consensus segmentation and Euclidean cluster extraction. |
|
Reconstruction |
Moving Least Squares (MLS) surface smoothing and Greedy Projection mesh meshing. |
3. Package Installation & Setup#
libpcl-oneapi packages are distributed through the Intel® Robotics AI Suite APT package repositories.
Supported Environments#
Ubuntu 24.04 LTS (Noble) / ROS 2 Jazzy
Ubuntu 22.04 LTS (Jammy) / ROS 2 Humble
Intel® oneAPI Toolkit: 2026.1 or newer
APT Repository Configuration#
# Add Intel oneAPI repository
curl -fsSL https://apt.repos.intel.com/intel-gpg-keys/GPG-PUB-KEY-INTEL-SW-PRODUCTS.PUB \
| sudo gpg --dearmor -o /usr/share/keyrings/oneapi-archive-keyring.gpg
echo "deb [signed-by=/usr/share/keyrings/oneapi-archive-keyring.gpg] https://apt.repos.intel.com/oneapi all main" \
| sudo tee /etc/apt/sources.list.d/oneAPI.list
# Add Intel Robotics AI Suite APT repository
echo "deb [trusted=yes] https://wheeljack.ch.intel.com/apt-repos/AMR/$(lsb_release -cs) amr main" \
| sudo tee /etc/apt/sources.list.d/intel-robotics-ai-suite.list
sudo apt-get update
Installing Packages#
Development Environment#
For building downstream robotics perception pipelines with oneAPI acceleration, install the compiler toolchain, peer FLANN acceleration library, and the libpcl-oneapi-dev development package (which automatically pulls libpcl-dev along with all required base and accelerated headers):
sudo apt-get install -y \
intel-oneapi-compiler-dpcpp-cpp-2026.1 \
intel-oneapi-libdpstd-devel \
intel-oneapi-tbb-devel-2023.1 \
libflann-oneapi-dev \
libpcl-oneapi-dev
sudo apt-get install -y \
intel-oneapi-compiler-dpcpp-cpp-2026.1 \
intel-oneapi-libdpstd-devel \
intel-oneapi-tbb-devel-2023.1 \
libflann-oneapi-dev \
libpcl-oneapi-dev
Production / Robot Deployment#
In deployment environments (such as onboard autonomous mobile robots or production container images) where code is executed without recompilation, install only the runtime libraries to minimize image footprint:
Install the full PCL acceleration runtime suite (including GPU drivers and runtime dependencies):
sudo apt-get install -y libpcl-oneapi
Install only the specific accelerated subsystems required by your application (for example, filtering and segmentation):
sudo apt-get install -y \
libpcl-oneapi-filters1.15 \
libpcl-oneapi-segmentation1.15
4. Consuming Application Integration Guide#
Consuming applications (such as ROS 2 perception nodes or standalone C++ robotics pipelines) integrate libpcl-oneapi using standard CMake target definitions.
Step 1: CMake Configuration#
In the downstream application CMakeLists.txt:
cmake_minimum_required(VERSION 3.25)
project(robotics_pointcloud_filter LANGUAGES CXX)
set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
# Find upstream PCL components
find_package(PCL CONFIG REQUIRED COMPONENTS common io)
# Find oneAPI PCL acceleration components
find_package(PCL-ONEAPI CONFIG REQUIRED COMPONENTS
oneapi_common
oneapi_filters
oneapi_features
oneapi_kdtree
oneapi_segmentation
)
add_executable(robotics_perception_node src/main.cpp)
# Enable SYCL compiler flags for Intel icpx
target_compile_options(robotics_perception_node PRIVATE
-fsycl
-fsycl-unnamed-lambda
-fsycl-device-code-split=per_kernel
)
target_link_options(robotics_perception_node PRIVATE -fsycl)
target_link_libraries(robotics_perception_node PRIVATE
${PCL_LIBRARIES}
${PCL-ONEAPI_LIBRARIES}
# Alternatively, link specific components:
# pcl_oneapi_common
# pcl_oneapi_filters
)
Step 2: Building the Downstream Application#
Important
Because SYCL acceleration utilizes device compilation models, applications must be compiled using Intel’s oneAPI C++ Compiler (icpx).
# 1. Initialize oneAPI environment
source /opt/intel/oneapi/setvars.sh
# 2. Configure with icpx
cmake -S . -B build -G Ninja \
-DCMAKE_CXX_COMPILER=icpx \
-DCMAKE_BUILD_TYPE=RelWithDebInfo
# 3. Build executable
cmake --build build --parallel
# 1. Initialize oneAPI and ROS environments
source /opt/intel/oneapi/setvars.sh
source /opt/ros/$ROS_DISTRO/setup.bash
# 2. Build ROS 2 package using icpx
colcon build \
--packages-select robotics_perception_node \
--cmake-args \
-DCMAKE_CXX_COMPILER=icpx \
-DCMAKE_BUILD_TYPE=RelWithDebInfo
Step 3: C++ Code Examples#
Device Queue Management: Automatic vs. Caller-Owned Queue#
Tip
libpcl-oneapi supports two execution models:
Automatic Device Management (Zero Boilerplate): Calling
.setQueue()is optional. When omitted, all algorithms automatically fall back topcl::oneapi::getDefaultQueue(). This default singleton queue queries the SYCL default selector and targets whatever device is configured in the environment (e.g., viaONEAPI_DEVICE_SELECTOR=level_zero:gpu). It also utilizes an internal device memory pool to minimize allocation overhead.Caller-Owned Queue Injection (
.setQueue(sycl::queue&)): Downstream applications can supply their ownsycl::queueto target specific GPU sub-devices, enforce in-order execution (sycl::property::queue::in_order), or interleave asynchronous compute with sensor acquisition. Note:PointCloudDev<T>instances and algorithms operating on them must share the samesycl::context.
Example 1: Accelerated Real-Time PassThrough & VoxelGrid Downsampling#
#include <iostream>
#include <memory>
#include <pcl/point_cloud.h>
#include <pcl/point_types.h>
#include <pcl/oneapi/filters/passthrough.h>
#include <pcl/oneapi/filters/voxel_grid.h>
int main()
{
auto input_cloud = std::make_shared<pcl::PointCloud<pcl::PointXYZ>>();
// ... populate input_cloud ...
// No SYCL boilerplate or queue management needed:
// Uses pcl::oneapi::getDefaultQueue() automatically.
auto z_filtered = std::make_shared<pcl::PointCloud<pcl::PointXYZ>>();
pcl::oneapi::PassThrough<pcl::PointXYZ> pass;
pass.setInputCloud(input_cloud);
pass.setFilterFieldName("z");
pass.setFilterLimits(0.2f, 1.8f);
pass.filter(*z_filtered);
auto voxel_filtered = std::make_shared<pcl::PointCloud<pcl::PointXYZ>>();
pcl::oneapi::VoxelGrid<pcl::PointXYZ> voxel;
voxel.setInputCloud(z_filtered);
voxel.setLeafSize(0.05f, 0.05f, 0.05f);
voxel.filter(*voxel_filtered);
return 0;
}
#include <iostream>
#include <memory>
#include <sycl/sycl.hpp>
#include <pcl/point_cloud.h>
#include <pcl/point_types.h>
#include <pcl/oneapi/filters/passthrough.h>
#include <pcl/oneapi/filters/voxel_grid.h>
int main()
{
// 1. Initialize an in-order GPU queue targeting the Intel GPU
sycl::queue q(sycl::gpu_selector_v, sycl::property::queue::in_order{});
std::cout << "Running on device: "
<< q.get_device().get_info<sycl::info::device::name>() << std::endl;
// 2. Allocate and generate input cloud (e.g. from 3D LiDAR)
auto input_cloud = std::make_shared<pcl::PointCloud<pcl::PointXYZ>>();
input_cloud->width = 100000;
input_cloud->height = 1;
input_cloud->points.resize(input_cloud->width * input_cloud->height);
for (size_t i = 0; i < input_cloud->size(); ++i) {
(*input_cloud)[i] = pcl::PointXYZ(
static_cast<float>(rand()) / RAND_MAX * 10.0f,
static_cast<float>(rand()) / RAND_MAX * 10.0f,
static_cast<float>(rand()) / RAND_MAX * 2.0f);
}
// 3. Stage 1: GPU-accelerated PassThrough Z-filtering with injected queue
auto z_filtered = std::make_shared<pcl::PointCloud<pcl::PointXYZ>>();
pcl::oneapi::PassThrough<pcl::PointXYZ> pass;
pass.setQueue(q);
pass.setInputCloud(input_cloud);
pass.setFilterFieldName("z");
pass.setFilterLimits(0.2f, 1.8f);
pass.filter(*z_filtered);
// 4. Stage 2: GPU-accelerated VoxelGrid Leaf Downsampling with injected queue
auto voxel_filtered = std::make_shared<pcl::PointCloud<pcl::PointXYZ>>();
pcl::oneapi::VoxelGrid<pcl::PointXYZ> voxel;
voxel.setQueue(q);
voxel.setInputCloud(z_filtered);
voxel.setLeafSize(0.05f, 0.05f, 0.05f);
voxel.filter(*voxel_filtered);
std::cout << "Original points: " << input_cloud->size()
<< " -> Downsampled points: " << voxel_filtered->size() << std::endl;
return 0;
}
Example 2: GPU Accelerated Surface Normal Estimation#
#include <sycl/sycl.hpp>
#include <pcl/point_cloud.h>
#include <pcl/point_types.h>
#include <pcl/oneapi/features/normal_3d.h>
#include <pcl/oneapi/search/kdtree.h>
void estimateNormals(const pcl::PointCloud<pcl::PointXYZ>::Ptr& cloud,
pcl::PointCloud<pcl::Normal>::Ptr& output_normals,
sycl::queue& q)
{
// Accelerated search tree backend
auto tree = std::make_shared<pcl::oneapi::search::KdTree<pcl::PointXYZ>>();
tree->setQueue(q);
// GPU Normal Estimation with warp reduction kernels
pcl::oneapi::NormalEstimation<pcl::PointXYZ, pcl::Normal> ne;
ne.setQueue(q);
ne.setInputCloud(cloud);
ne.setSearchMethod(tree);
ne.setRadiusSearch(0.15); // 15 cm neighborhood
ne.compute(*output_normals);
}
Example 3: RANSAC Ground Plane Segmentation#
#include <sycl/sycl.hpp>
#include <pcl/ModelCoefficients.h>
#include <pcl/PointIndices.h>
#include <pcl/oneapi/segmentation/sac_segmentation.h>
void extractGroundPlane(const pcl::PointCloud<pcl::PointXYZ>::Ptr& cloud,
pcl::PointIndices::Ptr& inliers,
pcl::ModelCoefficients::Ptr& coefficients,
sycl::queue& q)
{
pcl::oneapi::SACSegmentation<pcl::PointXYZ> seg;
seg.setQueue(q);
seg.setOptimizeCoefficients(true);
seg.setModelType(pcl::SACMODEL_PLANE);
seg.setMethodType(pcl::SAC_RANSAC);
seg.setMaxIterations(100);
seg.setDistanceThreshold(0.03); // 3 cm tolerance
seg.setInputCloud(cloud);
seg.segment(*inliers, *coefficients);
}
5. Runtime Device Selection and Execution#
Applications running libpcl-oneapi can dynamically select physical hardware using the oneAPI standard environment variables without modifying source code:
Target Device |
Runtime Variable Setting |
|---|---|
Intel Arc / Discrete GPU |
|
Intel Core Integrated GPU |
|
Intel Xeon CPU Fallback |
|
Specific Sub-Device / Tile |
|
Verification Command#
Verify that the runtime discovers the GPU compute devices correctly:
sycl-ls
Expected output:
[opencl:cpu] Intel(R) OpenCL, 13th Gen Intel(R) Core(TM) i9-13900K
[ext_oneapi_level_zero:gpu] Intel(R) Level-Zero, Intel(R) Arc(TM) A770 Graphics
6. Performance Best Practices#
Always initialize persistent sycl::queue instances at system startup with sycl::property::queue::in_order{}. Recreating queues per-frame introduces JIT context compilation and driver handshake latencies.
GPU offload is most efficient when clouds contain at least $15{,}000$ points. For sub-millimeter dense clusters or very small crops ($< 2{,}000$ points), standard CPU kernels or OpenMP dispatch should be benchmarked against GPU dispatch using the built-in PCL oneAPI benchmarks/ comparative harness.
On Intel Core Ultra and Xe processors, host CPU and integrated GPU share physical memory. Utilizing USM host/device shared allocations completely eliminates PCIe DMA transfer penalties.