Hang Yu, Xuebo Zhang, Zhenjie Zhao, Haochong Chen
IEEE TRANSACTIONS ON INDUSTRIAL ELECTRONICS
- The experimental video can be found here.
Generating reliable grasp configurations in cluttered scenes is an important guarantee for robots to successfully grasp objects. Existing learning-based methods encode the voxelized scene into a latent space to predict graspable points, which overlook both the reliability of grasp selection and the explicit geometric constraint imposed by the object's surface occupancy. In this paper, we propose a 3D sparse grasp network (3D SPGNet), which maps voxel features of object surfaces to grasp candidates to enforce explicit constraints, while leveraging the dense Truncated Signed Distance Function (TSDF) reconstruction to enhance grasp reliability through implicit synergy. Specifically, we first use a 3D CNN to extract multi-scale features at three different resolutions from the voxelized input. Then, we apply the proposed 3D surface constraint block to aggregate features and perform explicit geometric mapping on the encoded non-empty voxels. Finally, we reconstruct the scene from sparse to dense and generate the grasp configuration accordingly. Moreover, we propose a novel loss function for estimating reliable grasp scores and poses. Simulation experiments demonstrate that our method improves the grasp success rates and declutter rates by approximately 9% compared to state-of-the-art baselines. In addition, we deploy the model on a real robot arm, and the real-world experimental results show that our method achieves over a 10% improvement in performance, which further verifies the effectiveness of the proposed method. our results are shown in data/result.

-
Install the full version of CUDA Toolkit 11.1 (compatible with PyTorch and spconv-cu111).
-
Create a conda environment.
conda create -n spgnet3d python=3.8
- Activate the conda environment.
conda activate spgnet3d
- Install in the conda environment.
conda install pytorch-gpu=1.10.0 torchvision=0.9 cudatoolkit=11.1
- Install packages list in requirements.txt.
pip install -r requirements.txt -i https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple
- Then install
torch-scatterfollowing here, based onpytorchversion andcudaversion.
pip install torch-scatter -f https://data.pyg.org/whl/torch-1.10.0+cu111.html
- Go to the root directory and install the project locally using
pip.
pip install -e .
- Build ConvONets dependents by running
python scripts/convonet_setup.py build_ext --inplace
-
Install the AnyGrasp SDK. Note: to avoid package conflicts, graspnetAPI must be installed from source. Also, replace the
demo.pyfile in grasp_detection with ourdemo.py. -
We compare our method in the environment defined by GraspNeRF. You may treat GraspNeRF-main as an independent repository; installation and evaluation procedures follow those described in GraspNeRF-main/README or the original GraspNeRF repository. Note that torch_scatter in the
wheelsdirectory needs to be installed locally. The comparison results are shown inGraspNeRF-main/results. In addition, comparisons can also be conducted in our environment. You can download the GraspNeRF pretrained weights from here. After downloading, please place the file undersrc/nr/ckpt.
Pile scenario:
python scripts/generate_data_parallel.py --scene pile --object-set pile/train --num-grasps 16000000 --save-scene ./data/pile/data_pile_train_random_raw_16M --num-proc 1 --terminal-num 0 --grasps-per-scene 480Packed scenario:
python scripts/generate_data_parallel.py --scene packed --object-set packed/train --num-grasps 4000000 --save-scene ./data/pile/data_packed_train_random_raw_4M --num-proc 1 --terminal-num 0First clean and balance the data using: (pile)
python scripts/clean_balance_data.py ./data/pile/data_pile_train_random_raw_16M(packed)
python scripts/clean_balance_data.py ./data/pile/data_packed_train_random_raw_4MThen construct the dataset (add noise): (pile)
python scripts/construct_dataset_parallel.py --num-proc 1 --single-view --add-noise dex ./data/pile/data_pile_train_random_raw_16M ./data/new_dataset/data_pile_train_random_new_16M(packed)
python scripts/construct_dataset_parallel.py --num-proc 1 --single-view --add-noise dex ./data/pile/data_packed_train_random_raw_4M ./data/new_dataset/data_packed_train_random_new_4M(pile)
python scripts/save_occ_data_parallel.py ./data/pile/data_pile_train_random_raw_16M 100000 2 --num-proc 1(packed)
python scripts/save_occ_data_parallel.py ./data/pile/data_packed_train_random_raw_4M/ 100000 2 --num-proc 1Run: (pile)
python scripts/train.py --config config.yml --gpus 1 --scene pile --num 16(packed)
python scripts/train.py --config config.yml --gpus 1 --scene packed --num 4Run: (pile)
python scripts/sim_grasp_multiple.py --num-view 1 --object-set pile/test --scene pile --num-rounds 100 --sideview --add-noise dex --force --best --model data/models/spgrasp_pile.ckpt --type spg --result-path data/result/pile.json --config config.yml(packed)
python scripts/sim_grasp_multiple.py --num-view 1 --object-set packed/test --scene packed --num-rounds 100 --sideview --add-noise dex --force --best --model data/models/spgrasp_packed.ckpt --type spg --result-path data/result/packed.json --config config.ymlThis commands will run experiment with each seed specified in the arguments.
Pretrained models are in the data.zip. They are in data/models.
Data generation is very costly, so we upload the generated data . However, the occupancy data of GIGA exceeds 100 GB; therefore, we only uploaded the processed dataset. The new dataset alone is still sufficient to train our model and VGN. After downloading, extract it to the data folder under the repo's root.
| Raw data | Processed data |
|---|---|
| pile | new dataset |
@ARTICLE{11456265,
author={Yu, Hang and Zhang, Xuebo and Zhao, Zhenjie and Chen, Haochong},
journal={IEEE Transactions on Industrial Electronics},
title={3-D SPGNet: A 6-DoF Grasp Detection Network via 3-D Surface Constraint and TSDF Reconstruction},
year={2026},
volume={},
number={},
pages={1-13}
}