Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
Accelerator
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
Uh oh!
There was an error while loading.
Please reload this page
.
onnxruntime
/
mobius
Public
Notifications
You must be signed in to change notification settings
Fork
4
Star
17
Code
Issues
26
Pull requests
56
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Actions
Projects
Security and quality
Insights
Actions: onnxruntime/mobius
Actions
All workflows
Workflows
CI
CI
L4: Golden Checkpoint Parity (GPU)
L4: Golden Checkpoint Parity (GPU)
L5: End-to-End Generation (GPU)
L5: End-to-End Generation (GPU)
Nightly L2 Architecture Validation
Nightly L2 Architecture Validation
Architecture Diff
Architecture Diff
Benchmark
Benchmark
Check flags docs
Check flags docs
CodeQL
CodeQL
Copilot
Copilot
Copilot cloud agent
Copilot cloud agent
Show more workflows...
Management
Caches
Deployments
Benchmark
Benchmark
Actions
Loading...
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading.
Please reload this page
.
will be ignored since log searching is not yet available
Show workflow options
Create status badge
Create status badge
Loading
Uh oh!
There was an error while loading.
Please reload this page
.
benchmark.yml
will be ignored since log searching is not yet available
2,500+ workflow runs
2,500+ workflow runs
Event
Filter by Event
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching events.
Status
Filter by Status
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching statuses.
Branch
Filter by Branch
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching branches.
Actor
Filter by Actor
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching users.
fix(deepseek-v4): per-layer CSA/HCA compressed-record symbolic dim
Benchmark
#2650:
Pull request
#622
opened by
justinchuby
1m 51s
deckard/csa-per-layer-record-axis
deckard/csa-per-layer-record-axis
1m 51s
View #622
View workflow file
Fix real PLaMo2 cached decode parity
Benchmark
#2649:
Pull request
#620
synchronize by
justinchuby
1m 55s
justinchuby-fix-plamo2-runtime-parity
justinchuby-fix-plamo2-runtime-parity
1m 55s
View #620
View workflow file
Add dedicated Kimi-K3 GGUF support
Benchmark
#2648:
Pull request
#621
opened by
justinchuby
1m 48s
justinchuby-kimi-k3-gguf
justinchuby-kimi-k3-gguf
1m 48s
View #621
View workflow file
Fix real PLaMo2 cached decode parity
Benchmark
#2647:
Pull request
#620
opened by
justinchuby
1m 54s
justinchuby-fix-plamo2-runtime-parity
justinchuby-fix-plamo2-runtime-parity
1m 54s
View #620
View workflow file
Add dedicated Kimi Linear GGUF support
Benchmark
#2646:
Pull request
#619
opened by
justinchuby
1m 42s
justinchuby-kimi-linear-gguf
justinchuby-kimi-linear-gguf
1m 42s
View #619
View workflow file
Add ORT GenAI end-to-end CI
Benchmark
#2645:
Pull request
#618
opened by
justinchuby
1m 53s
justinchuby-ort-genai-e2e-ci
justinchuby-ort-genai-e2e-ci
1m 53s
View #618
View workflow file
Validate real PLaMo2 GGUF import
Benchmark
#2644:
Pull request
#617
opened by
justinchuby
1m 48s
justinchuby-validate-real-plamo2
justinchuby-validate-real-plamo2
1m 48s
View #617
View workflow file
Add exact MiniMax-01 GGUF support
Benchmark
#2643:
Pull request
#616
opened by
justinchuby
1m 57s
justinchuby-minimax-01-gguf
justinchuby-minimax-01-gguf
1m 57s
View #616
View workflow file
Add GraniteHybrid routed-MoE GGUF support
Benchmark
#2642:
Pull request
#615
opened by
justinchuby
1m 50s
justinchuby-granitehybrid-moe-gguf
justinchuby-granitehybrid-moe-gguf
1m 50s
View #615
View workflow file
Add Nemotron-H MoE GGUF support
Benchmark
#2641:
Pull request
#614
opened by
justinchuby
1m 58s
justinchuby-complete-nemotron-h-moe-gguf
justinchuby-complete-nemotron-h-moe-gguf
1m 58s
View #614
View workflow file
Add full Jamba GGUF MoE support
Benchmark
#2640:
Pull request
#613
opened by
justinchuby
1m 46s
justinchuby-complete-jamba-gguf
justinchuby-complete-jamba-gguf
1m 46s
View #613
View workflow file
Add dedicated PLaMo2 model and GGUF import
Benchmark
#2639:
Pull request
#612
opened by
justinchuby
1m 53s
justinchuby-plamo2-gguf
justinchuby-plamo2-gguf
1m 53s
View #612
View workflow file
Generate graph-driven ORT GenAI decoder configs
Benchmark
#2638:
Pull request
#611
opened by
justinchuby
1m 50s
justinchuby-generic-genai-configs
justinchuby-generic-genai-configs
1m 50s
View #611
View workflow file
Fail closed on lossy GGUF quantization preservation
Benchmark
#2637:
Pull request
#609
synchronize by
justinchuby
1m 49s
justinchuby-fix-gguf-q4-parity
justinchuby-fix-gguf-q4-parity
1m 49s
View #609
View workflow file
Add dedicated Falcon-H1 model and GGUF import
Benchmark
#2636:
Pull request
#610
opened by
justinchuby
1m 47s
justinchuby-falcon-h1-gguf
justinchuby-falcon-h1-gguf
1m 47s
View #610
View workflow file
Fail closed on lossy GGUF quantization preservation
Benchmark
#2635:
Pull request
#609
opened by
justinchuby
1m 54s
justinchuby-fix-gguf-q4-parity
justinchuby-fix-gguf-q4-parity
1m 54s
View #609
View workflow file
Materialize exact GGUF tokenizers
Benchmark
#2634:
Pull request
#608
opened by
justinchuby
1m 49s
justinchuby-materialize-gguf-tokenizers
justinchuby-materialize-gguf-tokenizers
1m 49s
View #608
View workflow file
Add small-model GGUF runtime evidence
Benchmark
#2633:
Pull request
#607
opened by
justinchuby
1m 51s
justinchuby-validate-small-gguf-models
justinchuby-validate-small-gguf-models
1m 51s
View #607
View workflow file
Add dedicated LFM2MoE GGUF graph support
Benchmark
#2632:
Pull request
#606
opened by
justinchuby
1m 46s
justinchuby-implement-deferred-recurrent-gguf
justinchuby-implement-deferred-recurrent-gguf
1m 46s
View #606
View workflow file
feat(deepseek-v4): default-off native CompressedSparseAttention (HCA ratio-128) export [C1, DRAFT]
Benchmark
#2631:
Pull request
#593
synchronize by
justinchuby
1m 46s
deckard/deepseek-v4-csa-hca-c1
deckard/deepseek-v4-csa-hca-c1
1m 46s
View #593
View workflow file
Finalize generated GGUF support matrix
Benchmark
#2630:
Pull request
#604
opened by
justinchuby
1m 44s
justinchuby-finalize-gguf-support-matrix
justinchuby-finalize-gguf-support-matrix
1m 44s
View #604
View workflow file
Add opt-in --paged-attention export for dense MLA (LATENT PagedAttention) [Slice 3B]
Benchmark
#2629:
Pull request
#599
synchronize by
justinchuby
1m 41s
squad/mobius-3b-paged-attention-export
squad/mobius-3b-paged-attention-export
1m 41s
View #599
View workflow file
Add opt-in --paged-attention export for dense MLA (LATENT PagedAttention) [Slice 3B]
Benchmark
#2628:
Pull request
#599
synchronize by
justinchuby
1m 48s
squad/mobius-3b-paged-attention-export
squad/mobius-3b-paged-attention-export
1m 48s
View #599
View workflow file
Fix Olive-quantized Qwen3 MoE export
Benchmark
#2627:
Pull request
#525
synchronize by
titaiwangms
1m 43s
fix/moe-packed-fused-expert-weights
fix/moe-packed-fused-expert-weights
1m 43s
View #525
View workflow file
Add opt-in --paged-attention export for dense MLA (LATENT PagedAttention) [Slice 3B]
Benchmark
#2626:
Pull request
#599
synchronize by
justinchuby
1m 39s
squad/mobius-3b-paged-attention-export
squad/mobius-3b-paged-attention-export
1m 39s
View #599
View workflow file
Previous
1
2
3
4
5
…
99
100
101
Next
You can’t perform that action at this time.