Trace every conclusion back to its evidence.
An open-source engine for reconstructing explainable behavioral sequences from heterogeneous security-analysis events while preserving an auditable chain from conclusions back to their original evidence.
🚧 Early development — research and engineering prototype
This project is currently in the architectural and foundational stage.
The core engine is not yet implemented.
The design may change as implementation, testing, research, and real-world evaluation reveal new requirements or limitations.
Modern malware-analysis and security-analysis tools generate large quantities of telemetry:
- process events;
- file operations;
- registry operations;
- memory operations;
- network activity;
- API calls;
- system calls;
- behavioral detections;
- ATT&CK mappings;
- sandbox observations.
These tools are extremely valuable, but there is an important gap between observing events and understanding the chain of behavior represented by those events.
An analyst may receive a result such as:
Process Injection detected
but still need to determine:
Why?
Which events caused this conclusion?
Which rule or reasoning process produced it?
Which events are directly observed?
Which relationships were inferred?
What evidence supports the conclusion?
Is there contradictory evidence?
Can the conclusion be traced back to the original event?
This project explores a different approach.
The goal is not to replace mature reverse-engineering, malware-analysis, or sandboxing platforms.
Instead, the project aims to provide a complementary reasoning layer capable of transforming heterogeneous observations into an explainable and auditable representation of behavior.
The central principle is:
Every conclusion should be traceable back to the evidence that supports it.
The project is built around a progression from observation to interpretation:
RAW EVENT
│
▼
EVIDENCE
│
▼
RELATION
│
▼
BEHAVIOR
│
▼
SEQUENCE
│
▼
INFERENCE
│
▼
CONCLUSION
The reverse path must also be possible:
CONCLUSION
│
▼
INFERENCE
│
▼
SEQUENCE
│
▼
BEHAVIOR
│
▼
RELATION
│
▼
EVIDENCE
│
▼
RAW EVENT
This bidirectional traceability is one of the fundamental design goals of the project.
The project is not trying to become another:
- Ghidra;
- IDA;
- Binary Ninja;
- radare2;
- x64dbg;
- malware sandbox;
- endpoint detection platform.
Instead, it focuses on a narrower problem:
How can heterogeneous security events be transformed into explainable behavioral sequences without losing the relationship between an interpretation and the evidence that produced it?
The intended architecture is complementary to existing analysis systems.
For example:
ANALYSIS SOURCES
│
┌────────────────┼────────────────┐
│ │ │
CAPE DRAKVUF capa
│ │ │
└────────────────┼────────────────┘
│
▼
CANONICAL EVENTS
│
▼
EVIDENCE MODEL
│
▼
CORRELATION ENGINE
│
▼
BEHAVIOR ENGINE
│
▼
SEQUENCE ENGINE
│
▼
INFERENCE ENGINE
│
▼
EXPLANATION ENGINE
│
▼
ANALYST
The system must distinguish between:
Observed fact
and:
Derived interpretation
For example:
Observed:
WriteProcessMemory()
is not equivalent to:
Process Injection
The second may be a behavioral interpretation requiring additional evidence and relationships.
A conclusion should be able to answer:
What evidence supports me?
Where did that evidence come from?
Which rule or reasoning process produced me?
Which events were used?
What relationships were inferred?
The engine should not only represent supporting evidence.
It should also be capable of representing evidence that weakens or contradicts an interpretation.
Example:
Conclusion:
Possible Persistence
Supporting evidence:
Scheduled Task created
Counter-evidence:
Scheduled Task immediately removed
Confidence:
Reduced
A confidence score should never be interpreted as absolute truth.
The system must preserve the distinction between:
Fact
Inference
Hypothesis
Explainability should not be added as a textual explanation after the analysis.
The underlying data model itself should preserve the relationships necessary to explain a conclusion.
The project should avoid unnecessarily creating proprietary representations when established standards are suitable.
Integration with concepts such as ATT&CK, STIX, and Attack Flow will be evaluated during development.
The initial conceptual model contains seven primary objects:
Event
Evidence
Relation
Behavior
Sequence
Inference
Conclusion
A directly observed occurrence.
Examples:
Process created
File written
OpenProcess called
TCP connection established
Registry key created
Information that establishes the provenance and context of an observation.
A relationship between observations or derived entities.
Examples:
precedes
causes
depends_on
same_process
same_target
supports
contradicts
A higher-level interpretation derived from one or more events and relationships.
A connected set of behaviors representing an observed or reconstructed progression.
An interpretation that goes beyond directly observed facts.
A resulting analytical statement containing its supporting reasoning and evidence.
Suppose the engine receives:
E01 OpenProcess()
E02 VirtualAllocEx()
E03 WriteProcessMemory()
E04 CreateRemoteThread()
The engine may construct:
E01 ─────┐
│
E02 ─────┤
├──> PROCESS_INJECTION
E03 ─────┤
│
E04 ─────┘
The behavioral conclusion must remain connected to the original observations:
PROCESS_INJECTION
│
├── Rule: PI-001
│
├── Evidence: E01
├── Evidence: E02
├── Evidence: E03
└── Evidence: E04
An analyst should be able to navigate:
Conclusion
↓
Behavior
↓
Evidence
↓
Event
↓
Original source
The initial scope focuses on:
- event representation;
- event normalization;
- evidence provenance;
- event correlation;
- behavioral reasoning;
- behavioral sequences;
- explainable inference;
- confidence;
- counter-evidence;
- graph-based relationships;
- interoperability with existing security-analysis ecosystems.
The project is not intended to initially become:
- a complete malware sandbox;
- a full disassembler;
- a debugger;
- an operating-system emulator;
- a replacement for Ghidra or IDA;
- an endpoint detection platform;
- an antivirus engine;
- an autonomous malware analyst.
Those capabilities may interact with the project but are not the initial objective.
This project is being developed as both an engineering project and a learning/research project.
Implementation decisions will be documented rather than hidden.
Important architectural decisions will be recorded as Architecture Decision Records (ADRs).
The development process will generally follow:
Problem
↓
Research
↓
Design
↓
Implementation
↓
Tests
↓
Documentation
↓
Review
↓
Commit
The core implementation is initially planned in:
C
The project may later incorporate:
- Assembly;
- platform-specific APIs;
- binary formats;
- security telemetry;
- graph algorithms;
- parsers;
- serialization formats;
- interoperability layers.
C was selected deliberately because the project provides an opportunity to study:
- memory;
- pointers;
- data structures;
- binary representation;
- compilation;
- debugging;
- low-level system behavior;
- the relationship between C and Assembly.
The project follows a gap-analysis approach.
The objective is not:
"Create something nobody has ever created."
Instead:
Find an existing capability, identify where it is incomplete, difficult, opaque, or poorly suited to a particular workflow, and build a significantly better solution for that specific problem.
Existing tools and approaches will be evaluated based on:
- What they do well.
- What they do poorly.
- What they do not do.
- What is unnecessarily complex.
- What is difficult for beginners.
- What remains difficult for experienced analysts.
- What information is hidden.
- What information is poorly represented.
- What workflow steps require manual work.
- What recurring user complaints exist.
- What issues or discussions reveal.
- What use cases remain underserved.
Claims about gaps should be supported by evidence whenever possible.
The project is designed to investigate interoperability and complementarity with technologies such as:
- CAPE Sandbox;
- DRAKVUF;
- capa;
- MITRE ATT&CK;
- MITRE Attack Flow;
- GReAT-BeD;
- ANY.RUN;
- Joe Sandbox.
These projects are not treated as competitors by default.
The purpose of studying them is to understand the current ecosystem and identify a defensible gap.
| Component | Status |
|---|---|
| Project definition | 🟢 Initial |
| Research | 🟡 Ongoing |
| Domain model | 🟡 Designing |
| Event model | 🔴 Not implemented |
| Evidence model | 🔴 Not implemented |
| Relation model | 🔴 Not implemented |
| Behavior engine | 🔴 Not implemented |
| Sequence engine | 🔴 Not implemented |
| Inference engine | 🔴 Not implemented |
| External adapters | 🔴 Not implemented |
| Tests | 🔴 Not implemented |
| CLI | 🔴 Not implemented |
See ROADMAP.md.
Architecture documentation:
Research:
Architecture decisions:
See CONTRIBUTING.md.
See SECURITY.md.
This project is licensed under the MIT License.
See LICENSE.
This project is intended for legitimate security research, malware analysis, education, defensive security engineering, and interoperability research.
It is not intended to facilitate unauthorized access, compromise, persistence, or deployment of malware against systems without authorization.