Performance Results
Measured performance for the hbac-rs, abac-rs, and ldap-acis evaluation engines
across rule counts from 100 to 1,000,000. All numbers are single-threaded,
release-mode, on synthetic workloads using the sssd-prod distribution preset.
Test Environments
ARM64 — Apple M4
- Platform: macOS 15.4 (ARM64)
- CPU: Apple M4 (10 cores, single-threaded benchmarks)
- Rust: 1.96.0
- Build:
cargo build --release(opt-level=3)
x86_64 — Intel Core i7-12800H
- Platform: Linux 6.15.8 (Fedora 42)
- CPU: 12th Gen Intel Core i7-12800H (14 cores, single-threaded benchmarks)
- Build:
cargo build --release(opt-level=3)
Performance is platform-dependent. Always benchmark on your target hardware.
Note: These results reflect the latest optimizations. See the Cross-BAC Benchmarking guide for detailed optimization journey and cross-engine comparisons.
All engines use AHash (fast non-cryptographic hashing) for all internal hash operations.
HBAC Performance
Cached (LRU-warm)
10,000 matching requests after 1,000 warmup requests. The LRU cache dominates.
ARM64 (Apple M4)
| Rules | Throughput | Mean | P95 | P99 | Memory |
|---|---|---|---|---|---|
| 100 | 7.85M r/s | 127 ns | 208 ns | 292 ns | 34 KB |
| 1,000 | 5.49M r/s | 182 ns | 292 ns | 375 ns | 334 KB |
| 10,000 | 4.07M r/s | 245 ns | 375 ns | 417 ns | 3.3 MB |
| 20,000 | 3.36M r/s | 298 ns | 459 ns | 542 ns | 6.6 MB |
| 50,000 | 1.68M r/s | 596 ns | 1.00 µs | 1.21 µs | 16.5 MB |
| 100,000 | 947K r/s | 1.06 µs | 1.83 µs | 2.21 µs | 33 MB |
| 250,000 | 428K r/s | 2.34 µs | 4.17 µs | 5.00 µs | 82.5 MB |
| 500,000 | 215K r/s | 4.64 µs | 8.33 µs | 9.96 µs | 165 MB |
| 1,000,000 | 109K r/s | 9.14 µs | 16.5 µs | 19.8 µs | 330 MB |
x86_64 (Intel i7-12800H)
| Rules | Throughput | Mean | P95 | P99 | Memory |
|---|---|---|---|---|---|
| 100 | 16.0M r/s | 62 ns | 71 ns | 77 ns | 34 KB |
| 1,000 | 15.1M r/s | 66 ns | 75 ns | 82 ns | 334 KB |
| 10,000 | 13.9M r/s | 72 ns | 81 ns | 88 ns | 3.2 MB |
| 20,000 | 13.9M r/s | 72 ns | 81 ns | 87 ns | 6.4 MB |
| 50,000 | 14.1M r/s | 71 ns | 81 ns | 86 ns | 16.1 MB |
| 100,000 | 11.5M r/s | 87 ns | 99 ns | 106 ns | 32.3 MB |
| 250,000 | 11.7M r/s | 86 ns | 99 ns | 106 ns | 80.7 MB |
| 500,000 | 7.74M r/s | 129 ns | 151 ns | 166 ns | 161.5 MB |
| 1,000,000 | 11.3M r/s | 88 ns | 102 ns | 109 ns | 323.1 MB |
check_access (LRU-warm)
check_access() returns bool instead of HbacEvaluationResult, avoiding
the 3×Vec<String> allocation overhead.
x86_64 (Intel i7-12800H)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 19.2M r/s | 52 ns | 60 ns | 66 ns |
| 1,000 | 16.8M r/s | 60 ns | 70 ns | 78 ns |
| 10,000 | 15.9M r/s | 63 ns | 73 ns | 78 ns |
| 20,000 | 16.6M r/s | 60 ns | 70 ns | 75 ns |
| 50,000 | 13.8M r/s | 73 ns | 86 ns | 91 ns |
| 100,000 | 13.2M r/s | 76 ns | 88 ns | 95 ns |
| 250,000 | 13.6M r/s | 74 ns | 84 ns | 90 ns |
| 500,000 | 13.0M r/s | 77 ns | 90 ns | 98 ns |
| 1,000,000 | 12.7M r/s | 79 ns | 92 ns | 101 ns |
Uncached
10,000 unique non-matching requests, 0 warmup. Every request exercises the full evaluation path (Bloom filter, decision tree, deny-rule index). The first request at 10K+ triggers an index rebuild, creating a bimodal distribution (P95/P99 reflect steady-state; the mean includes the cold first request).
ARM64 (Apple M4)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 11.0M r/s | 91 ns | 166 ns | 167 ns |
| 1,000 | 8.15M r/s | 123 ns | 167 ns | 208 ns |
| 10,000 | 6.40M r/s | 156 ns | 125 ns | 167 ns |
| 20,000 | 3.88M r/s | 258 ns | 166 ns | 167 ns |
| 50,000 | 2.03M r/s | 492 ns | 125 ns | 167 ns |
| 100,000 | 1.06M r/s | 941 ns | 167 ns | 209 ns |
| 250,000 | 449K r/s | 2.23 µs | 250 ns | 292 ns |
| 500,000 | 223K r/s | 4.48 µs | 375 ns | 417 ns |
| 1,000,000 | 106K r/s | 9.42 µs | 625 ns | 708 ns |
x86_64 (Intel i7-12800H)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 11.2M r/s | 89 ns | 118 ns | 174 ns |
| 1,000 | 6.50M r/s | 154 ns | 195 ns | 237 ns |
| 10,000 | 2.50M r/s | 400 ns | 302 ns | 379 ns |
| 20,000 | 1.27M r/s | 787 ns | 374 ns | 463 ns |
| 50,000 | 373K r/s | 2.68 µs | 908 ns | 1.21 µs |
| 100,000 | 249K r/s | 4.02 µs | 684 ns | 828 ns |
| 250,000 | 90K r/s | 11.1 µs | 786 ns | 946 ns |
| 500,000 | 47K r/s | 21.3 µs | 971 ns | 1.15 µs |
| 1,000,000 | 23K r/s | 43.7 µs | 1.25 µs | 1.44 µs |
Mixed Workload (Throughput)
80% matching + 20% non-matching, cache enabled, 1,000 warmup.
ARM64 (Apple M4)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 9.26M r/s | 108 ns | 167 ns | 250 ns |
| 1,000 | 6.37M r/s | 157 ns | 250 ns | 292 ns |
| 10,000 | 5.55M r/s | 180 ns | 292 ns | 334 ns |
| 20,000 | 3.65M r/s | 274 ns | 459 ns | 542 ns |
| 50,000 | 1.91M r/s | 523 ns | 1.00 µs | 1.21 µs |
| 100,000 | 1.05M r/s | 955 ns | 1.96 µs | 2.33 µs |
| 250,000 | 485K r/s | 2.06 µs | 4.33 µs | 5.21 µs |
| 500,000 | 249K r/s | 4.02 µs | 8.46 µs | 10.2 µs |
| 1,000,000 | 129K r/s | 7.74 µs | 16.4 µs | 19.7 µs |
x86_64 (Intel i7-12800H)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 15.6M r/s | 64 ns | 73 ns | 78 ns |
| 1,000 | 14.8M r/s | 68 ns | 79 ns | 86 ns |
| 10,000 | 11.3M r/s | 88 ns | 101 ns | 111 ns |
| 20,000 | 11.7M r/s | 86 ns | 103 ns | 112 ns |
| 50,000 | 11.0M r/s | 91 ns | 108 ns | 122 ns |
| 100,000 | 9.58M r/s | 104 ns | 125 ns | 140 ns |
| 250,000 | 9.31M r/s | 107 ns | 130 ns | 189 ns |
| 500,000 | 9.87M r/s | 101 ns | 120 ns | 143 ns |
| 1,000,000 | 9.17M r/s | 109 ns | 131 ns | 176 ns |
ABAC Performance
With AHash optimization and bitmap-based deny indexing, ABAC delivers peak throughput at 1K–10K rules and scales inversely with deny rule count beyond that. On x86_64, throughput ranges from 5.7M r/s at 1K rules to 143K r/s at 1M rules, while using 4.2× less memory than HBAC. HBAC now uses the same bitmap deny index optimization and is faster at all scales, but ABAC provides N-dimensional flexibility.
Cached (Single Request Latency)
ARM64 (Apple M4)
| Rules | Throughput | Mean | P95 | P99 | Memory |
|---|---|---|---|---|---|
| 100 | 1.48M r/s | 0.67µs | 1.46µs | 2.17µs | 7.8 KB |
| 1,000 | 3.19M r/s | 0.31µs | 0.42µs | 0.50µs | 78 KB |
| 10,000 | 3.73M r/s | 0.27µs | 0.38µs | 0.46µs | 781 KB |
| 20,000 | 3.06M r/s | 0.33µs | 0.50µs | 0.62µs | 1.6 MB |
| 50,000 | 1.70M r/s | 0.59µs | 1.04µs | 1.29µs | 3.9 MB |
| 100,000 | 969K r/s | 1.03µs | 1.92µs | 2.42µs | 7.8 MB |
| 250,000 | 435K r/s | 2.30µs | 4.38µs | 5.38µs | 19.5 MB |
| 500,000 | 223K r/s | 4.48µs | 8.67µs | 10.9µs | 39 MB |
| 1,000,000 | 113K r/s | 8.84µs | 17.2µs | 21.3µs | 78 MB |
x86_64 (Intel i7-12800H)
| Rules | Throughput | Mean | P95 | P99 | Memory |
|---|---|---|---|---|---|
| 100 | 2.22M r/s | 450 ns | 939 ns | 1.14 µs | 7.8 KB |
| 1,000 | 5.69M r/s | 176 ns | 227 ns | 247 ns | 78 KB |
| 10,000 | 3.97M r/s | 252 ns | 348 ns | 398 ns | 781 KB |
| 20,000 | 2.88M r/s | 348 ns | 559 ns | 710 ns | 1.5 MB |
| 50,000 | 1.67M r/s | 598 ns | 974 ns | 1.19 µs | 3.8 MB |
| 100,000 | 949K r/s | 1.05 µs | 1.73 µs | 2.15 µs | 7.6 MB |
| 250,000 | 546K r/s | 1.83 µs | 3.18 µs | 4.24 µs | 19.1 MB |
| 500,000 | 289K r/s | 3.46 µs | 6.16 µs | 9.22 µs | 38.1 MB |
| 1,000,000 | 143K r/s | 7.01 µs | 12.9 µs | 15.4 µs | 76.3 MB |
Uncached
10,000 unique non-matching requests, 0 warmup. The first request triggers compiled evaluator construction, dominating the mean latency. P95/P99 reflect steady-state uncached performance.
ARM64 (Apple M4)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 930K r/s | 1.08µs | 1.21µs | 1.33µs |
| 1,000 | 3.78M r/s | 0.26µs | 0.21µs | 0.25µs |
| 10,000 | 1.40M r/s | 0.72µs | 0.17µs | 0.21µs |
| 20,000 | 671K r/s | 1.49µs | 0.17µs | 0.25µs |
| 50,000 | 269K r/s | 3.72µs | 0.17µs | 0.25µs |
| 100,000 | 161K r/s | 6.19µs | 0.21µs | 0.25µs |
| 250,000 | 57.8K r/s | 17.3µs | 0.21µs | 0.25µs |
| 500,000 | 27.8K r/s | 36.0µs | 0.33µs | 0.38µs |
| 1,000,000 | 13.3K r/s | 75.1µs | 0.50µs | 0.54µs |
x86_64 (Intel i7-12800H)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 786K r/s | 1.27 µs | 1.40 µs | 1.64 µs |
| 1,000 | 3.92M r/s | 255 ns | 235 ns | 292 ns |
| 10,000 | 836K r/s | 1.20 µs | 596 ns | 739 ns |
| 20,000 | 346K r/s | 2.89 µs | 785 ns | 966 ns |
| 50,000 | 160K r/s | 6.26 µs | 865 ns | 1.00 µs |
| 100,000 | 80K r/s | 12.5 µs | 944 ns | 1.10 µs |
| 250,000 | 25K r/s | 39.3 µs | 1.68 µs | 1.89 µs |
| 500,000 | 16K r/s | 61.7 µs | 1.25 µs | 1.39 µs |
| 1,000,000 | 7.6K r/s | 131 µs | 1.71 µs | 1.88 µs |
Mixed Workload (Throughput)
80% matching + 20% non-matching, cache enabled, 1,000 warmup.
ARM64 (Apple M4)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 1.93M r/s | 518 ns | 1.00 µs | 1.17 µs |
| 1,000 | 5.21M r/s | 192 ns | 250 ns | 292 ns |
| 10,000 | 4.18M r/s | 239 ns | 375 ns | 417 ns |
| 20,000 | 3.20M r/s | 313 ns | 542 ns | 625 ns |
| 50,000 | 1.88M r/s | 531 ns | 1.04 µs | 1.29 µs |
| 100,000 | 1.09M r/s | 919 ns | 1.96 µs | 2.46 µs |
| 250,000 | 500K r/s | 2.00 µs | 4.50 µs | 5.67 µs |
| 500,000 | 259K r/s | 3.86 µs | 8.79 µs | 11.0 µs |
| 1,000,000 | 131K r/s | 7.61 µs | 17.6 µs | 22.2 µs |
x86_64 (Intel i7-12800H)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 1.58M r/s | 631 ns | 1.31 µs | 1.43 µs |
| 1,000 | 5.98M r/s | 167 ns | 218 ns | 246 ns |
| 10,000 | 3.48M r/s | 287 ns | 427 ns | 520 ns |
| 20,000 | 2.85M r/s | 351 ns | 537 ns | 626 ns |
| 50,000 | 1.55M r/s | 646 ns | 1.13 µs | 1.36 µs |
| 100,000 | 718K r/s | 1.39 µs | 2.62 µs | 3.16 µs |
| 250,000 | 524K r/s | 1.91 µs | 3.87 µs | 4.83 µs |
| 500,000 | 348K r/s | 2.87 µs | 6.33 µs | 7.88 µs |
| 1,000,000 | 192K r/s | 5.20 µs | 11.2 µs | 14.0 µs |
LDAP ACI Performance
LDAP Access Control Instructions use a DnScopeIndex trie for O(depth)
candidate selection, with dual LRU caches (check_access + authorize) in
CachedAciPolicy. Cached evaluation is essentially O(1) — latency stays
flat at 47–75 ns regardless of rule count. Uncached evaluation uses
objectclass pre-filtering, userdn pre-filtering, and attribute partitioning
to eliminate ~75% of candidates before per-ACI checks, scaling from 598 ns
at 100 rules to 1.5 ms at 1M rules.
The synthetic workload matches real FreeIPA ACI distributions: 60% permission-based (GroupDn + ObjectClass filter + targetattr), 15% authenticated read (ObjectClass filter), 10% admin group, 5% anonymous read, 5% self-service, 5% narrow UserDn scope. Non-matching requests use valid DNs with mismatched objectclasses, exercising the full filter/bind/attr rejection pipeline rather than trivially failing at the DN trie.
Cached (LRU-warm)
10,000 matching requests after 1,000 warmup requests. The LRU cache dominates.
x86_64 (Intel i7-12800H)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 16.5M r/s | 60 ns | 66 ns | 76 ns |
| 1,000 | 14.3M r/s | 70 ns | 76 ns | 86 ns |
| 10,000 | 14.5M r/s | 69 ns | 75 ns | 86 ns |
| 20,000 | 14.4M r/s | 69 ns | 77 ns | 85 ns |
| 50,000 | 15.1M r/s | 66 ns | 74 ns | 81 ns |
| 100,000 | 15.0M r/s | 66 ns | 75 ns | 83 ns |
| 250,000 | 15.2M r/s | 66 ns | 71 ns | 81 ns |
| 500,000 | 16.0M r/s | 62 ns | 70 ns | 75 ns |
| 1,000,000 | 13.4M r/s | 75 ns | 83 ns | 93 ns |
check_access (LRU-warm)
check_access() returns bool instead of AuthorizationResult, avoiding
the PermissionSet allocation overhead.
x86_64 (Intel i7-12800H)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 21.1M r/s | 47 ns | 53 ns | 58 ns |
| 1,000 | 18.0M r/s | 55 ns | 61 ns | 65 ns |
| 10,000 | 18.5M r/s | 54 ns | 61 ns | 64 ns |
| 20,000 | 18.3M r/s | 55 ns | 61 ns | 65 ns |
| 50,000 | 19.7M r/s | 51 ns | 58 ns | 63 ns |
| 100,000 | 19.1M r/s | 52 ns | 61 ns | 67 ns |
| 250,000 | 19.5M r/s | 51 ns | 58 ns | 63 ns |
| 500,000 | 17.3M r/s | 58 ns | 66 ns | 69 ns |
| 1,000,000 | 16.5M r/s | 61 ns | 68 ns | 72 ns |
Uncached
10,000 unique non-matching requests, 0 warmup. Every request exercises the full evaluation path (DN trie traversal, objectclass/userdn pre-filtering, scope matching, bind rule checks). Non-matching requests use valid DNs with mismatched objectclasses (“device”), so they pass the DN trie but are rejected by filter/bind/attr checks — measuring realistic rejection cost.
x86_64 (Intel i7-12800H)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 1.7M r/s | 598 ns | 734 ns | 833 ns |
| 1,000 | 667K r/s | 1.50 µs | 1.68 µs | 1.78 µs |
| 10,000 | 205K r/s | 4.89 µs | 5.53 µs | 9.37 µs |
| 20,000 | 105K r/s | 9.53 µs | 10.91 µs | 15.46 µs |
| 50,000 | 40K r/s | 25.2 µs | 28.4 µs | 32.6 µs |
| 100,000 | 15K r/s | 65.2 µs | 89.4 µs | 104 µs |
| 250,000 | 5K r/s | 221 µs | 280 µs | 299 µs |
| 500,000 | 2K r/s | 587 µs | 714 µs | 832 µs |
| 1,000,000 | 1K r/s | 1.50 ms | 1.86 ms | 1.98 ms |
Mixed Workload (Throughput)
80% matching + 20% non-matching, cache enabled, 1,000 warmup.
x86_64 (Intel i7-12800H)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 15.1M r/s | 66 ns | 73 ns | 85 ns |
| 1,000 | 14.1M r/s | 71 ns | 81 ns | 92 ns |
| 10,000 | 12.7M r/s | 79 ns | 90 ns | 99 ns |
| 20,000 | 11.3M r/s | 89 ns | 102 ns | 114 ns |
| 50,000 | 11.5M r/s | 87 ns | 102 ns | 168 ns |
| 100,000 | 12.7M r/s | 78 ns | 89 ns | 104 ns |
| 250,000 | 12.3M r/s | 81 ns | 90 ns | 154 ns |
| 500,000 | 12.1M r/s | 83 ns | 96 ns | 174 ns |
| 1,000,000 | 13.6M r/s | 74 ns | 86 ns | 157 ns |
Windows SD Performance
Access check latency for Windows Security Descriptors using the check_access
algorithm (DACL walk with deny-before-allow, SID matching, owner implicit
rights). The round-robin adapter evaluates one SD per request, so latency
reflects the per-SD access check cost.
Unlike HBAC/ABAC (which search across a rule set), a Windows access check walks a single SD’s DACL — typically 3–8 ACEs. Latency is therefore nearly constant regardless of how many SDs are loaded; the rule count affects only memory and load time.
ARM64 (Apple M4)
Check Access Latency
| Rules | Mean (ns) | P95 (ns) | P99 (ns) | Throughput (req/s) |
|---|---|---|---|---|
| 100 | 44 | 83 | 84 | 22.6M |
| 1,000 | 38 | 42 | 42 | 26.4M |
| 10,000 | 23 | 42 | 84 | 43.1M |
| 100,000 | 21 | 42 | 125 | 48.7M |
Throughput (mixed matching/non-matching)
| Rules | Mean (ns) | P95 (ns) | Throughput (req/s) |
|---|---|---|---|
| 100 | 59 | 84 | 17.1M |
| 1,000 | 39 | 42 | 25.4M |
| 10,000 | 23 | 42 | 42.9M |
| 100,000 | 21 | 42 | 47.9M |
Uncached Latency (diverse non-matching)
| Rules | Mean (ns) | P95 (ns) | Throughput (req/s) |
|---|---|---|---|
| 100 | 49 | 83 | 20.5M |
| 1,000 | 36 | 42 | 27.5M |
| 10,000 | 22 | 42 | 45.3M |
| 100,000 | 21 | 42 | 47.5M |
Policy Build Time and Memory
| Rules | Build (ms) | Memory (KB) | Per rule (B) |
|---|---|---|---|
| 100 | 0.04 | 13.3 | 136 |
| 1,000 | 0.20 | 132.8 | 136 |
| 10,000 | 1.09 | 1328.1 | 136 |
| 100,000 | 9.40 | 13281.3 | 136 |
Cross-engine comparison (check_access, 10K rules)
| Engine | Mean (ns) | P95 (ns) | Throughput (req/s) |
|---|---|---|---|
| win-sd | 23 | 42 | 43.1M |
| posix-acl | 29 | 42 | 34.4M |
Windows SD access checks are slightly faster than POSIX ACL checks despite the more complex permission model (32-bit masks vs 3-bit rwx, explicit deny ACEs, SID matching vs string comparison). Both are in the same performance class at sub-50ns latency.
Policy Build Time
One-time cost when a policy is loaded or reloaded. Includes Bloom filter construction, decision tree compilation, deny-rule index building, and composite index creation for HBAC, or compiled evaluator, bitmap deny index, and composite index for ABAC, or DN trie index construction and LRU cache initialization for LDAP ACI, or simple Vec loading for Windows SD and POSIX ACL.
HBAC Build Time
ARM64 (Apple M4)
| Rules | Build Time | Per 1K Rules |
|---|---|---|
| 100 | 0.01 ms | 0.12 ms |
| 1,000 | 0.09 ms | 0.09 ms |
| 10,000 | 0.71 ms | 0.07 ms |
| 20,000 | 1.52 ms | 0.08 ms |
| 50,000 | 3.81 ms | 0.08 ms |
| 100,000 | 8.02 ms | 0.08 ms |
| 250,000 | 20.8 ms | 0.08 ms |
| 500,000 | 40.9 ms | 0.08 ms |
| 1,000,000 | 87.6 ms | 0.09 ms |
x86_64 (Intel i7-12800H)
| Rules | Build Time | Per 1K Rules |
|---|---|---|
| 100 | 0.04 ms | 0.38 ms |
| 1,000 | 0.13 ms | 0.13 ms |
| 10,000 | 1.57 ms | 0.16 ms |
| 20,000 | 5.57 ms | 0.28 ms |
| 50,000 | 19.4 ms | 0.39 ms |
| 100,000 | 41.7 ms | 0.42 ms |
| 250,000 | 106 ms | 0.43 ms |
| 500,000 | 202 ms | 0.40 ms |
| 1,000,000 | 424 ms | 0.42 ms |
ABAC Build Time
ARM64 (Apple M4)
| Rules | Build Time | Per 1K Rules |
|---|---|---|
| 100 | 0.15 ms | 1.50 ms |
| 1,000 | 0.83 ms | 0.83 ms |
| 10,000 | 6.19 ms | 0.62 ms |
| 20,000 | 13.4 ms | 0.67 ms |
| 50,000 | 33.4 ms | 0.67 ms |
| 100,000 | 60.7 ms | 0.61 ms |
| 250,000 | 167 ms | 0.67 ms |
| 500,000 | 354 ms | 0.71 ms |
| 1,000,000 | 798 ms | 0.80 ms |
x86_64 (Intel i7-12800H)
| Rules | Build Time | Per 1K Rules |
|---|---|---|
| 100 | 0.11 ms | 1.10 ms |
| 1,000 | 0.47 ms | 0.47 ms |
| 10,000 | 8.79 ms | 0.88 ms |
| 20,000 | 19.8 ms | 0.99 ms |
| 50,000 | 56.6 ms | 1.13 ms |
| 100,000 | 122 ms | 1.22 ms |
| 250,000 | 317 ms | 1.27 ms |
| 500,000 | 611 ms | 1.22 ms |
| 1,000,000 | 1,175 ms | 1.18 ms |
LDAP ACI Build Time
Build time includes DN trie construction, objectclass/userdn pre-filtering index building, and attribute partitioning.
x86_64 (Intel i7-12800H)
| Rules | Build Time | Per 1K Rules |
|---|---|---|
| 100 | 0.26 ms | 2.58 ms |
| 1,000 | 1.76 ms | 1.76 ms |
| 10,000 | 20.1 ms | 2.01 ms |
| 20,000 | 38.3 ms | 1.92 ms |
| 50,000 | 104 ms | 2.07 ms |
| 100,000 | 199 ms | 1.99 ms |
| 250,000 | 530 ms | 2.12 ms |
| 500,000 | 1.04 s | 2.09 ms |
| 1,000,000 | 2.1 s | 2.06 ms |
Memory
HBAC Memory Usage
ARM64 (Apple M4)
| Rules | Total | Per-Rule |
|---|---|---|
| 100 | 34 KB | 0.35 KB |
| 1,000 | 334 KB | 0.34 KB |
| 10,000 | 3.3 MB | 0.34 KB |
| 20,000 | 6.6 MB | 0.34 KB |
| 50,000 | 16.5 MB | 0.34 KB |
| 100,000 | 33 MB | 0.34 KB |
| 250,000 | 82.5 MB | 0.34 KB |
| 500,000 | 165 MB | 0.34 KB |
| 1,000,000 | 330 MB | 0.34 KB |
x86_64 (Intel i7-12800H)
| Rules | Total | Per-Rule |
|---|---|---|
| 100 | 34 KB | 0.34 KB |
| 1,000 | 334 KB | 0.34 KB |
| 10,000 | 3.2 MB | 0.33 KB |
| 20,000 | 6.4 MB | 0.33 KB |
| 50,000 | 16.1 MB | 0.33 KB |
| 100,000 | 32.3 MB | 0.33 KB |
| 250,000 | 80.7 MB | 0.33 KB |
| 500,000 | 161.5 MB | 0.33 KB |
| 1,000,000 | 323.1 MB | 0.33 KB |
ABAC Memory Usage
ARM64 (Apple M4)
| Rules | Total | Per-Rule |
|---|---|---|
| 100 | 7.8 KB | 0.080 KB |
| 1,000 | 78 KB | 0.080 KB |
| 10,000 | 781 KB | 0.080 KB |
| 20,000 | 1.6 MB | 0.080 KB |
| 50,000 | 3.9 MB | 0.080 KB |
| 100,000 | 7.8 MB | 0.080 KB |
| 250,000 | 19.5 MB | 0.080 KB |
| 500,000 | 39 MB | 0.080 KB |
| 1,000,000 | 78 MB | 0.080 KB |
x86_64 (Intel i7-12800H)
| Rules | Total | Per-Rule |
|---|---|---|
| 100 | 7.8 KB | 0.078 KB |
| 1,000 | 78 KB | 0.078 KB |
| 10,000 | 781 KB | 0.078 KB |
| 20,000 | 1.5 MB | 0.078 KB |
| 50,000 | 3.8 MB | 0.078 KB |
| 100,000 | 7.6 MB | 0.078 KB |
| 250,000 | 19 MB | 0.078 KB |
| 500,000 | 38 MB | 0.078 KB |
| 1,000,000 | 76 MB | 0.078 KB |
ABAC uses 4.2× less memory than HBAC due to the compiled evaluator avoiding per-rule storage overhead.
Per-rule cost is the average for the sssd-prod distribution. Category=all
rules are smaller (no entity strings); specific rules with many entities are
larger.
Index Overhead
- LRU cache: configurable (default 1,024 entries) × ~64 bytes = ~64 KB
- Bloom filters: ~1 KB per 1,000 rules
- Decision trees: ~0.05 KB per deny rule
- Deny-rule index: ~1 KB per deny rule (bitmap indexes over user/host/service dimensions)
- Total: <1% of rule storage at 10K+ rules
Performance Crossover
HBAC vs ABAC vs LDAP ACI — Mixed Workload (x86_64)
| Rules | HBAC | ABAC | LDAP ACI | Fastest |
|---|---|---|---|---|
| 100 | 15.6M r/s | 1.58M r/s | 15.1M r/s | HBAC |
| 1,000 | 14.8M r/s | 5.98M r/s | 14.1M r/s | HBAC |
| 10,000 | 11.3M r/s | 3.48M r/s | 12.7M r/s | LDAP ACI |
| 20,000 | 11.7M r/s | 2.85M r/s | 11.3M r/s | HBAC |
| 50,000 | 11.0M r/s | 1.55M r/s | 11.5M r/s | LDAP ACI |
| 100,000 | 9.58M r/s | 718K r/s | 12.7M r/s | LDAP ACI |
| 250,000 | 9.31M r/s | 524K r/s | 12.3M r/s | LDAP ACI |
| 500,000 | 9.87M r/s | 348K r/s | 12.1M r/s | LDAP ACI |
| 1,000,000 | 9.17M r/s | 192K r/s | 13.6M r/s | LDAP ACI |
LDAP ACI is the fastest engine at 10K+ rules in the mixed workload, with
throughput of 11–15M r/s thanks to its dual LRU cache and DnScopeIndex
trie. HBAC is slightly faster at small scales (100–1K rules) due to lower
per-request overhead, while ABAC lags further due to rule traversal even on
cache hits. The mixed workload uses realistic non-matching requests that
exercise objectclass/userdn pre-filtering and attribute rejection, so the
20% non-matching portion has slightly higher latency than trivial rejections.
Cached Workloads (typical SSSD deployments)
ARM64 (Apple M4):
- At 100 rules: ABAC is 2.7× faster (7.41M vs 2.78M req/s)
- At 100K rules: ABAC is 3.3× faster (830K vs 252K req/s)
x86_64 (Intel i7-12800H):
- At 100 rules: HBAC is 7.2× faster than ABAC (16.0M vs 2.22M req/s)
- At 1K rules: HBAC fastest (14.7M r/s) > LDAP ACI (14.3M) > ABAC (5.69M)
- At 100K rules: LDAP ACI (15.0M) > HBAC (11.5M) > ABAC (949K r/s)
- At 1M rules: LDAP ACI (13.4M) > HBAC (11.3M) > ABAC (143K r/s)
Uncached Workloads
ARM64 (Apple M4):
- At 10K rules: HBAC is 6.6× faster uncached (1.18M vs 179K req/s)
- At 100K rules: HBAC is 8.6× faster uncached (108K vs 12.6K req/s)
x86_64 (Intel i7-12800H):
- At 1K rules: HBAC (6.50M) > ABAC (3.92M) > LDAP ACI (667K r/s)
- At 10K rules: HBAC (2.50M) > ABAC (836K) > LDAP ACI (205K r/s)
- At 100K rules: HBAC (249K) > ABAC (80K) > LDAP ACI (15K r/s)
- At 1M rules: HBAC (23K) > ABAC (7.6K) > LDAP ACI (1K r/s)
HBAC and ABAC are faster than LDAP ACI for uncached workloads. LDAP ACI uncached numbers reflect a realistic workload where non-matching requests use valid DNs that pass the DN trie but are rejected by objectclass pre-filtering, userdn pre-filtering, and attribute checks — measuring the true cost of filter/bind rejection rather than trivial DN misses. HBAC and ABAC use Bloom filters and decision trees for faster uncached rejection.
Recommendation: For cached workloads, LDAP ACI is the fastest engine. For uncached workloads, HBAC is dramatically faster. Use ABAC when you need N-dimensional attribute-based rules or 4.2× less memory. Choose the engine that matches your access control model.
Optimization Layers
ABAC uses a 6-layer optimization pipeline:
- Constant-result fast path – single outcome (~15 ns)
- Bitmap deny index – u64 bitmask intersection when universal allow exists
- LRU cache – memoization (sub-microsecond, for non-indexed fallback)
- AHash – 2-3× faster than SipHash for non-cryptographic use
- Compiled evaluator – pre-extracted attributes + array indexing
- Composite index – candidate selection fallback (O(log n))
- Deny-only indexing – skip allow rules when universal allow exists
Key optimization: The bitmap deny index (Layer 1) uses pre-computed BitPos structs and per-dimension inverted indexes with u64 bitmask AND operations. When a universal allow rule exists, evaluation bypasses the cache entirely and resolves via bitmap intersection in O(deny_rules/64) per dimension. This produces excellent throughput at 1K–10K rules (3–5M r/s) where the bitmaps fit in L1 cache, with graceful degradation at larger scales.
Layer 4 (Compiled evaluator) pre-extracts request attributes once (3 HashMap lookups) into a stack-allocated array, then uses array indexing for all rule checks. This serves as the fallback when no universal allow rule exists.
Temporal Rule Support
Both HBAC and ABAC support time-based rules via TemporalHbacRule and
TemporalAbacRule. Temporal rules are evaluated separately and incur minimal
overhead when no temporal rules are active (~5-10ns for the empty-check).
When temporal rules are present:
- Active temporal rules (within time window) are evaluated alongside regular rules
- Inactive rules (outside time window) are skipped at evaluation time
- No additional caching overhead
- Build time increases by ~1-2ms per 1,000 temporal rules
See the Temporal Rules guide for usage examples.
JIT Compilation
Optional feature flag (--features jit). Compiles per-rule evaluation to
native code.
| Workload | Effect |
|---|---|
| Cached | No measurable difference (±3-6% noise). The LRU cache already provides sub-microsecond lookups. |
| Uncached | ~8-12% throughput improvement at all scales. Every request exercises the full evaluation path where native code pays off. |
| Build | No consistent overhead (±8% noise). |
JIT is worth enabling when uncached or low-hit-rate workloads dominate. For typical SSSD deployments with high cache hit rates, it provides no benefit.
Optimization Stack
Evaluations pass through these layers in order. Each layer is skipped if the previous one already produced a result.
Layer 0 — Constant-Result Fast Path
When the decision tree collapses to a single outcome (e.g., a universal deny rule exists, or a universal allow with no deny rules), the result is stored and returned immediately. Bypasses all other layers. ~15 ns per evaluation.
Layer 1 — Bitmap Deny Index
When a universal allow rule exists, the deny index short-circuits evaluation using bitmap intersection. Pre-computed BitPos structs map each deny rule to a (word, mask) pair. Per-dimension inverted indexes map attribute values to their rule bitmasks. Evaluation ANDs dimension bitmasks together; any surviving bit means a deny match. Cost: O(deny_rules/64) per dimension. Bypasses the LRU cache entirely — no RequestKey hashing overhead.
Layer 2 — LRU Memoization Cache
Configurable-size cache (default 1,024 entries) keyed by a u64 hash of (user, host, service, groups). Order-independent group hashing via wrapping addition. Zero heap allocations per lookup. Sub-100ns hit latency. Used when the bitmap deny index is not applicable.
Layer 3 — Bloom Filter Pre-screening
Fast negative check before full evaluation. 1% false positive rate. Effective
for non-matching requests when no category=all rules exist. In SSSD
deployments where category=all is common, the Bloom filter rarely rejects.
Layer 4 — Decision Tree
Rules are compiled into a decision tree at load time. When a universal allow exists alongside specific deny rules, the tree stores only the deny rules — reducing evaluation from O(n) to O(d) where d is the deny rule count.
Layer 5 — Composite Indexing (Fallback)
Indexed user+host lookup for candidate selection. Used when the decision tree
is not available. O(1) for category=all, O(log n) for specific entities.
Layer 6 — JIT Compilation (optional)
Requires the jit feature flag. Compiles per-rule evaluation to native code.
~10% throughput gain on uncached workloads; no gain on cached workloads.
Deployment Guidance
Typical SSSD (100–1,000 rules)
Most FreeIPA environments fall here.
HBAC (x86_64): 62–66 ns cached latency (evaluate), 52–60 ns (check_access), 15.1–16.0M req/s throughput, 34–334 KB memory, <0.13 ms build time.
ABAC (x86_64): 176–450 ns cached latency, 2.22–5.69M req/s throughput, 7.8–78 KB memory, <0.47 ms build time.
LDAP ACI (x86_64): 60–70 ns cached latency, 47–55 ns (check_access), 14.3–16.5M req/s, <1.76 ms build time.
Recommendation: LDAP ACI and HBAC deliver near-identical cached latency at this scale. Use ABAC when you need N-dimensional flexibility or 4.2× less memory. Choose the engine that matches your access control model.
Large Enterprise (1K–10K rules)
HBAC (x86_64): 66–72 ns cached latency, 13.9–15.1M req/s, 334 KB–3.2 MB memory, 0.13–1.57 ms build time.
ABAC (x86_64): 176–252 ns cached latency, 3.97–5.69M req/s, 78–781 KB memory, 0.47–8.79 ms build time.
LDAP ACI (x86_64): 69–70 ns cached latency, 54–55 ns (check_access), 14.3–14.5M req/s, 1.76–20.1 ms build time.
Recommendation: HBAC and LDAP ACI are close at this scale (14.3M vs 14.8M r/s at 1K rules). Both use LRU caching for flat latency. Use ABAC for N-dimensional flexibility or lower memory (4.2× less).
Very Large (10K–100K rules)
HBAC (x86_64): 72–87 ns cached latency, 11.5–13.9M req/s, 3.2–32.3 MB memory, 1.57–41.7 ms build time.
ABAC (x86_64): 252 ns–1.05 µs cached latency, 949K–3.97M req/s, 781 KB–7.6 MB memory, 8.79–122 ms build time.
LDAP ACI (x86_64): 66–69 ns cached latency, 51–52 ns (check_access), 14.4–15.1M req/s, 20.1–199 ms build time.
LDAP ACI is 1.3× faster than HBAC at 100K rules (15.0M vs 11.5M r/s). ABAC uses 4.2× less memory than HBAC.
Recommendation: LDAP ACI and HBAC are both excellent at this scale. Choose ABAC if memory is a constraint (7.6 MB vs 32.3 MB at 100K rules) or you need N-dimensional rules.
Extreme Scale (100K–1M rules)
HBAC (x86_64): 87–88 ns cached latency, 11.3–11.5M req/s, 32.3–323 MB memory, 41.7–424 ms build time.
ABAC (x86_64): 1.05–7.01 µs cached latency, 143K–949K req/s, 7.6–76 MB memory, 122 ms–1.18 s build time.
LDAP ACI (x86_64): 62–75 ns cached latency, 51–61 ns (check_access), 13.4–16.0M req/s, 199 ms–2.1 s build time.
At 1M rules, LDAP ACI is 1.2× faster than HBAC (13.4M vs 11.3M r/s) and 94× faster than ABAC (13.4M vs 143K r/s).
Recommendation: LDAP ACI delivers the best cached throughput at extreme scale. HBAC is close behind. Choose ABAC if memory is the primary constraint (76 MB vs 323 MB at 1M rules) or you need N-dimensional rules.
Cache Tuning
The default 1,024-entry LRU is optimal for most workloads.
- Higher hit rate: 2,048 or 4,096 entries (+64-192 KB memory)
- Memory constrained: 512 entries (~85% hit rate)
- Cold-start sensitive: pre-warm with common requests after policy load
- Diverse patterns: increase cache size if hit rate drops below 80%
Benchmark Methodology
Scenarios
| Scenario | Requests | Warmup | Distribution |
|---|---|---|---|
single-latency | 10,000 | 1,000 | All matching (pool: 200) |
check-access-latency | 10,000 | 1,000 | All matching (pool: 200, bool) |
uncached-latency | 10,000 | 0 | All non-matching (pool: 5,000) |
throughput | 10,000 | 1,000 | 80/20 mixed (pool: 500) |
build-time | 1 | 0 | Matching |
All evaluation scenarios pre-generate a request pool before measurement.
Requests cycle through the pool during warmup and measurement, ensuring
that only evaluation time is measured — not request object construction.
check-access-latency uses check_access() (returns bool) instead of
evaluate() (returns full result with matched rule lists).
Fixtures
Synthetic fixtures use the sssd-prod distribution, which models real SSSD
deployments: 15% user category=all, 30% host category=all, 40% service
category=all, 5% deny rules, 2% disabled rules. Universal deny rules (all
three dimensions category=all with type=deny) are excluded — real deployments
never create these, and they would collapse the decision tree to a constant.
LDAP ACI fixtures use the ldap-aci distribution, which models real 389-ds
deployments with varied DN scopes, bind rules (identity + environment), target
filters, attribute sets, and grant/deny ACIs.
Measurement
- Seeded RNG:
StdRng::seed_from_u64(42)for reproducibility - Pre-generated pools: requests are created before timing begins
- Single-threaded: sequential request processing
- Throughput: derived from the sum of per-request
evaluate()latencies - Timing: only the
evaluate()/check_access()call is measured
Running Benchmarks
Individual fixture generation and benchmarking:
# Generate fixtures
cargo run --release -p perf-testing -- \
generate -b hbac -c 10000 -o hbac_10000_baseline --fixtures-dir fixtures
# Run all scenarios for HBAC
cargo run --release -p perf-testing bench -b hbac -f hbac_10000_baseline --cache
# Run all scenarios for ABAC
cargo run --release -p perf-testing bench -b abac -f abac_10000_baseline --cache
# Run all scenarios for LDAP ACI
cargo run --release -p perf-testing bench -b ldap_aci -f ldap_aci_10000_baseline --cache
# Run specific scenarios
cargo run --release -p perf-testing -- \
list # List available fixtures and scenarios
Full benchmark suite across all scales:
# HBAC benchmarks
scripts/bench.sh run --no-jit --rules 100,1000,10000,20000,50000,100000,250000,500000,1000000
# ABAC benchmarks
scripts/bench.sh run --no-jit --bac-type abac --rules 100,1000,10000,20000,50000,100000,250000,500000,1000000
# LDAP ACI benchmarks
scripts/bench.sh run --no-jit --bac-type ldap_aci --rules 100,1000,10000,20000,50000,100000,250000,500000,1000000
Criterion Benchmarks
For statistically rigorous measurements with confidence intervals:
cd crates/perf-testing
cargo bench --bench hbac_benchmarks
open target/criterion/report/index.html
Python Binding Benchmarks
For the Python binding performance results below:
source .venv/bin/activate
python python/collect_results.py --warmup 3 --measure 5
Python Binding Performance
Measured on x86_64 — Intel Core i7-12800H (Linux 6.15.8, Python 3.15). Same fixtures and request-generation logic as the Rust benchmarks. Policy is built once per fixture and shared across all eval scenarios.
Two build/eval strategies are benchmarked:
- evaluate+builder: constructs rules via PyO3 builder pattern, evaluates
with
evaluate()(returns fullDecision/HbacEvaluationResult) - check_access+json (HBAC only): loads rules via
load_rules_json(), evaluates withcheck_access()(returnsbool, avoidsVec<String>allocation per call)
For ABAC, load_rules_json() provides faster build times (no Python object
construction) with identical evaluation speed — both paths produce the same
internal rules.
Python — HBAC (evaluate+builder)
Cached (LRU-warm, matching requests)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 6.43M r/s | 155 ns | 162 ns | 167 ns |
| 1,000 | 6.39M r/s | 157 ns | 161 ns | 165 ns |
| 10,000 | 6.26M r/s | 160 ns | 167 ns | 168 ns |
| 20,000 | 6.53M r/s | 153 ns | 161 ns | 165 ns |
| 50,000 | 6.46M r/s | 155 ns | 159 ns | 160 ns |
| 100,000 | 6.53M r/s | 153 ns | 156 ns | 163 ns |
| 250,000 | 6.38M r/s | 157 ns | 166 ns | 170 ns |
| 500,000 | 6.38M r/s | 157 ns | 161 ns | 164 ns |
| 1,000,000 | 6.21M r/s | 161 ns | 168 ns | 169 ns |
Uncached (~80% cache misses, non-matching)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 5.85M r/s | 171 ns | 183 ns | 186 ns |
| 1,000 | 4.39M r/s | 228 ns | 239 ns | 240 ns |
| 10,000 | 4.69M r/s | 213 ns | 227 ns | 233 ns |
| 20,000 | 4.18M r/s | 239 ns | 243 ns | 251 ns |
| 50,000 | 4.15M r/s | 241 ns | 246 ns | 249 ns |
| 100,000 | 3.65M r/s | 274 ns | 284 ns | 289 ns |
| 250,000 | 2.88M r/s | 347 ns | 366 ns | 371 ns |
| 500,000 | 2.02M r/s | 496 ns | 507 ns | 519 ns |
| 1,000,000 | 1.37M r/s | 731 ns | 749 ns | 754 ns |
Mixed Workload (80% matching / 20% non-matching)
| Rules | Throughput | Mean |
|---|---|---|
| 100 | 6.47M r/s | 154 ns |
| 1,000 | 6.18M r/s | 162 ns |
| 10,000 | 6.46M r/s | 155 ns |
| 20,000 | 6.35M r/s | 157 ns |
| 50,000 | 6.32M r/s | 158 ns |
| 100,000 | 6.39M r/s | 156 ns |
| 250,000 | 6.32M r/s | 158 ns |
| 500,000 | 6.23M r/s | 161 ns |
| 1,000,000 | 6.01M r/s | 166 ns |
Policy Build Time
| Rules | Rule objects (py) | load_rules() (rs) | Total |
|---|---|---|---|
| 100 | 0.7 ms | 0.0 ms | 0.7 ms |
| 1,000 | 2.2 ms | 0.1 ms | 2.3 ms |
| 10,000 | 20.1 ms | 1.6 ms | 21.7 ms |
| 20,000 | 36.8 ms | 3.0 ms | 39.8 ms |
| 50,000 | 99.3 ms | 8.1 ms | 107.4 ms |
| 100,000 | 186.8 ms | 15.0 ms | 201.8 ms |
| 250,000 | 475.3 ms | 36.3 ms | 511.5 ms |
| 500,000 | 981.2 ms | 70.0 ms | 1,051 ms |
| 1,000,000 | 1,938 ms | 142.4 ms | 2,080 ms |
Python — HBAC (check_access+json)
The optimized path: load_rules_json() skips Python rule-object allocation
entirely, and check_access() returns a bool instead of allocating
HbacEvaluationResult (3×Vec<String>).
Cached (LRU-warm, matching requests)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 7.93M r/s | 126 ns | 131 ns | 134 ns |
| 1,000 | 8.25M r/s | 121 ns | 126 ns | 127 ns |
| 10,000 | 8.25M r/s | 121 ns | 126 ns | 128 ns |
| 20,000 | 8.15M r/s | 123 ns | 126 ns | 128 ns |
| 50,000 | 8.16M r/s | 123 ns | 126 ns | 127 ns |
| 100,000 | 8.14M r/s | 123 ns | 127 ns | 131 ns |
| 250,000 | 7.71M r/s | 130 ns | 138 ns | 159 ns |
| 500,000 | 7.68M r/s | 130 ns | 136 ns | 140 ns |
| 1,000,000 | 8.15M r/s | 123 ns | 127 ns | 130 ns |
Uncached (~80% cache misses, non-matching)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 6.46M r/s | 155 ns | 165 ns | 175 ns |
| 1,000 | 5.41M r/s | 185 ns | 203 ns | 207 ns |
| 10,000 | 5.53M r/s | 181 ns | 189 ns | 195 ns |
| 20,000 | 5.25M r/s | 191 ns | 196 ns | 205 ns |
| 50,000 | 4.87M r/s | 205 ns | 211 ns | 212 ns |
| 100,000 | 4.10M r/s | 244 ns | 258 ns | 261 ns |
| 250,000 | 3.07M r/s | 326 ns | 349 ns | 360 ns |
| 500,000 | 2.21M r/s | 452 ns | 471 ns | 480 ns |
| 1,000,000 | 1.35M r/s | 739 ns | 820 ns | 859 ns |
Mixed Workload (80% matching / 20% non-matching)
| Rules | Throughput | Mean |
|---|---|---|
| 100 | 7.54M r/s | 133 ns |
| 1,000 | 7.75M r/s | 129 ns |
| 10,000 | 8.09M r/s | 124 ns |
| 20,000 | 8.10M r/s | 123 ns |
| 50,000 | 7.93M r/s | 126 ns |
| 100,000 | 7.77M r/s | 129 ns |
| 250,000 | 7.46M r/s | 134 ns |
| 500,000 | 7.82M r/s | 128 ns |
| 1,000,000 | 7.62M r/s | 131 ns |
Policy Build Time (JSON)
| Rules | load_rules_json() (rs) | Speedup vs builder |
|---|---|---|
| 100 | 0.6 ms | 1.2× |
| 1,000 | 1.2 ms | 1.9× |
| 10,000 | 6.9 ms | 3.1× |
| 20,000 | 12.8 ms | 3.1× |
| 50,000 | 32.0 ms | 3.4× |
| 100,000 | 63.2 ms | 3.2× |
| 250,000 | 160.6 ms | 3.2× |
| 500,000 | 313.9 ms | 3.3× |
| 1,000,000 | 617.4 ms | 3.4× |
Python — ABAC (evaluate+builder)
Cached (LRU-warm, matching requests)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 1.66M r/s | 603 ns | 621 ns | 626 ns |
| 1,000 | 3.12M r/s | 320 ns | 330 ns | 334 ns |
| 10,000 | 2.44M r/s | 410 ns | 422 ns | 426 ns |
| 20,000 | 2.22M r/s | 451 ns | 467 ns | 478 ns |
| 50,000 | 1.54M r/s | 649 ns | 683 ns | 723 ns |
| 100,000 | 1.02M r/s | 982 ns | 1.04 µs | 1.14 µs |
| 250,000 | 524K r/s | 1.91 µs | 2.05 µs | 2.07 µs |
| 500,000 | 278K r/s | 3.60 µs | 3.69 µs | 3.78 µs |
| 1,000,000 | 159K r/s | 6.29 µs | 7.13 µs | 7.18 µs |
ABAC evaluation scales inversely with rule count — no FFI floor like the prior (incorrect) ARM64 measurements suggested. At 1K rules the compiled evaluator delivers 3.1M r/s; at 1M rules the larger index scan drops to 159K.
Uncached (~80% cache misses, non-matching)
| Rules | Throughput | Mean | P95 | P99 |
|---|---|---|---|---|
| 100 | 649K r/s | 1.54 µs | 1.66 µs | 1.69 µs |
| 1,000 | 3.41M r/s | 293 ns | 310 ns | 316 ns |
| 10,000 | 3.35M r/s | 299 ns | 316 ns | 322 ns |
| 20,000 | 3.46M r/s | 289 ns | 298 ns | 304 ns |
| 50,000 | 3.32M r/s | 301 ns | 310 ns | 316 ns |
| 100,000 | 3.13M r/s | 319 ns | 328 ns | 336 ns |
| 250,000 | 2.56M r/s | 390 ns | 404 ns | 420 ns |
| 500,000 | 1.94M r/s | 515 ns | 534 ns | 562 ns |
| 1,000,000 | 1.55M r/s | 645 ns | 671 ns | 691 ns |
Uncached is faster than cached at 1K+ rules because the Bloom filter rejects non-matching requests before rule iteration.
Mixed Workload (80% matching / 20% non-matching)
| Rules | Throughput | Mean |
|---|---|---|
| 100 | 1.12M r/s | 896 ns |
| 1,000 | 3.17M r/s | 315 ns |
| 10,000 | 2.57M r/s | 389 ns |
| 20,000 | 2.28M r/s | 439 ns |
| 50,000 | 1.66M r/s | 603 ns |
| 100,000 | 1.20M r/s | 835 ns |
| 250,000 | 633K r/s | 1.58 µs |
| 500,000 | 337K r/s | 2.97 µs |
| 1,000,000 | 198K r/s | 5.06 µs |
Policy Build Time
| Rules | Rule objects (py) | load_rules() (rs) | Total |
|---|---|---|---|
| 100 | 0.7 ms | 0.1 ms | 0.7 ms |
| 1,000 | 3.3 ms | 0.7 ms | 3.9 ms |
| 10,000 | 30.5 ms | 7.5 ms | 38.0 ms |
| 20,000 | 59.8 ms | 15.0 ms | 74.8 ms |
| 50,000 | 148.1 ms | 34.0 ms | 182.1 ms |
| 100,000 | 304.1 ms | 68.1 ms | 372.2 ms |
| 250,000 | 788.5 ms | 183.2 ms | 971.7 ms |
| 500,000 | 1,564 ms | 370.3 ms | 1,935 ms |
| 1,000,000 | 3,123 ms | 769.5 ms | 3,892 ms |
Using load_rules_json() instead of the builder pattern is 2.6–3.1× faster
for build time (same eval performance):
| Rules | Builder Total | JSON Total | Speedup |
|---|---|---|---|
| 1,000 | 3.9 ms | 1.5 ms | 2.6× |
| 10,000 | 38.0 ms | 13.3 ms | 2.9× |
| 100,000 | 372.2 ms | 121.3 ms | 3.1× |
| 1,000,000 | 3,892 ms | 1,354 ms | 2.9× |
Python vs Rust — Summary (x86_64)
| Metric | Rust (x86_64) | Python (x86_64) | Ratio |
|---|---|---|---|
| HBAC evaluate (cached, 1K) | 66 ns | 157 ns | Rust 2.4× faster |
| HBAC check_access (cached, 1K) | 60 ns | 121 ns | Rust 2.0× faster |
| HBAC evaluate (cached, 100K) | 87 ns | 153 ns | Rust 1.8× faster |
| HBAC check_access (cached, 100K) | 76 ns | 123 ns | Rust 1.6× faster |
| HBAC throughput (1K) | 14.8M r/s | 6.18M r/s | Rust 2.4× faster |
| HBAC build (10K, builder) | 1.57 ms | 21.7 ms | ~14× slower |
| HBAC build (10K, json) | 1.57 ms | 6.9 ms | ~4.4× slower |
| ABAC eval (cached, 1K) | 176 ns | 320 ns | Rust 1.8× faster |
| ABAC eval (cached, 100K) | 1.05 µs | 982 ns | ~equal |
| ABAC throughput (10K) | 3.48M r/s | 2.57M r/s | Rust 1.4× faster |
| ABAC build (10K, builder) | 8.79 ms | 38.0 ms | ~4.3× slower |
| ABAC build (10K, json) | 8.79 ms | 13.3 ms | ~1.5× slower |
Key observations:
-
HBAC evaluation: With pre-generated request pools (fixing the old benchmark methodology that counted request allocation), native Rust is 1.6–2.4× faster than Python. The
check_access()path (returningboolinstead ofHbacEvaluationResult) is the fastest option in both Rust and Python. Python’s#[pyclass(frozen)]optimization still delivers impressive throughput (8.25M r/s at 1K rules). -
ABAC evaluation: At small scales (1K), Rust is ~1.8× faster. At 100K rules, Python is on par — the compiled evaluator dominates and the PyO3 call overhead is negligible relative to rule evaluation time.
-
Build time: The
load_rules_json()path eliminates Python rule-object construction entirely, achieving build times within 1.5–4.4× of native Rust. The builder path is 4–14× slower due to per-rule FFI round-trips. -
Use case: Python bindings deliver near-native evaluation performance and are suitable for production data-plane use. Use
load_rules_json()andcheck_access()(HBAC) for maximum throughput.