This is an official repository for ProvG-Searcher: A Graph Representation Learning Approach for Efficient Provenance Graph Search.
The repository provides three main capabilities for using ProvG-Searcher:
- Calculating Sampling Stats: It calculates the sampling statistics used during training.
- Training the Model: It trains the model on Darpha datasets.
- Testing the Model: It tests the models on Darpha datasets.
You can set up the conda environment with all the requirements using the following command:
conda env create -f environment.ymlFor Darpha datasets, ego graphs and node features can be obtained from Drive. Unzip the files and place them in the data and node_feature folders, respectively.
The model can be trained for each dataset as follows:
python -u run.py --dataset data/[dataset_name]/k_[number_of_neighborhood].pt --data_identifier [dataset_name] --model_path ckpt/[dataset_name]_model.pthHere, [dataset_name] can be one of:
- ta1-theia-e3-official-6r
- fiveDirection
- cadets_data
- trace_data
and [number_of_neighborhood] can be 3 or 5.
To test a model, set the testing arguments to True.
If you use this code in your work, please cite the accompanying paper:
@inproceedings{10.1145/3576915.3623187,
author = {Altinisik, Enes and Deniz, Fatih and Sencar, H\"{u}srev Taha},
title = {ProvG-Searcher: A Graph Representation Learning Approach for Efficient Provenance Graph Search},
year = {2023},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3576915.3623187},
doi = {10.1145/3576915.3623187},
booktitle = {Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security},
pages = {2247–2261},
numpages = {15},
}
The model part of the code is forked from neural-subgraph-learning-GNN repository.