This repository contains pre-built llama.cpp binaries optimized for the NVIDIA Jetson Orin platform.
- Build date: May 10, 2025
- Target platform: NVIDIA Jetson Orin
- CUDA version: 12.x
- llama.cpp version: Latest from main branch
bin/llama-cli- Main command-line interface for text generationbin/llama-server- HTTP/WebSocket server for inferencebin/llama-quantize- Tool for quantizing models to smaller sizesbin/llama-perplexity- Measure perplexity of a model on textbin/llama-tokenize- Convert text to and from token IDs
bin/llama-bench- Benchmarking tool for performance testingbin/llama-gguf- Utility for inspecting and modifying GGUF filesbin/llama-simple- Minimal implementation for basic text generationbin/llama-batched- Implementation supporting batched processingbin/llama-speculative- Speculative decoding implementation
bin/llama-llava-cli- CLI for vision-language models (LLAVA)bin/llama-mtmd-cli- Multimodal model interfacebin/llama-gemma3-cli- Interface for Gemma 3 modelsbin/llama-qwen2vl-cli- Interface for Qwen2-VL models
bin/llama-gguf-split- Tool for splitting large GGUF filesbin/llama-gguf-hash- Tool for hashing GGUF filesbin/llama-embedding- Generate embeddings from text
lib/libllama.so- Core llama.cpp shared librarylib/libggml.so- GGML tensor librarylib/libggml-base.so- GGML base componentslib/libggml-cpu.so- CPU-specific implementationslib/libllava_shared.so- LLAVA shared componentslib/libmtmd_shared.so- Multimodal shared components
./bin/llama-cli -m /path/to/model.gguf -p "Write a short poem about AI:" -n 200