Component: val/src/rule_based_orchestrator.c (sysarch-acs)
Type: Enhancement / runtime optimization
Severity: Low correctness risk, high runtime impact
Summary
The rule-based orchestrator re-executes a child rule once for every parent rule
that references it. Results are already recorded in rule_status_map[], but the
map is never consulted before execution, so identical work is repeated.
On our platform this roughly triples ACS wall time. In a full run that completed
at 555.0 ms of simulated time, 328.4 ms of the 506.1 ms rule span (65%) was
spent on second-and-later dispatches of rules that had already produced a
result.
Evidence
Worst offender is PCI_IN_02 ("Check ECAM Memory accessibility"), which walks
the entire ECAM (buses 0x00-0xFF x 32 dev x 8 func):
PCI_IN_16 ("Check all 1's for out of range") behaves identically: four
dispatches of 31.2 ms each, 124.9 ms total, 93.7 ms of it redundant.
Together these two rules account for 341.5 ms of the 555.0 ms suite (62%).
Aggregate cost of repeat dispatches (2nd and later), same run:
| ms |
dispatches |
rule |
| 163.5 |
4 |
PCI_IN_02 |
| 93.7 |
4 |
PCI_IN_16 |
| 27.8 |
3 |
PCI_SM_02 |
| 16.4 |
3 |
PCI_IN_05 |
| 3.6 |
3 |
PCI_LI_01 |
| 2.9 |
3 |
PCI_MSI_2 |
| 2.8 |
3 |
PCI_IN_19 |
| 2.7 |
3 |
PCI_LI_03 |
Repetition is widespread, not limited to PCIe. Of 372 distinct rules observed:
285 dispatched once, 21 twice, 58 three times, and 8 four times.
Root cause
In execute_rule_recursive():
- Before descending, the function guards only against cycles
(rule_reference_path_contains()) and explicit skips (is_rule_skipped()).
There is no check for "this rule already produced a result in this run".
- Child rules are dispatched unconditionally:
child_rule_status = execute_rule_recursive(ctx, child_rule_id, 1, num_pe, 1);
- On exit, the result is stored, for children as well as top-level rules:
if (report_self) {
rule_status_map[rule_id] = rule_test_status;
print_rule_test_status(rule_id, indent, rule_test_status);
}
rule_status_map[] is therefore already populated with exactly the value a
cache would need. It is currently used only for reporting.
The result of a BASE_RULE is a pure function of the rule and the PE count -
the dispatch carries no parent-specific context:
rule_test_status = test_entry_func_table[rule_test_map[rule_id].test_entry_id](num_pe);
This is what makes reuse sound in principle: a second execution under a
different parent is given identical inputs, so it can only produce an identical
result (absent side effects - see Risks).
Proposed enhancement
Consult rule_status_map[] before executing a rule, and reuse a recorded
result instead of re-running the test.
Suggested shape:
- Add an explicit "not yet executed" sentinel and clear
rule_status_map[] at
the start of each run (it must not persist across runs).
- Early in
execute_rule_recursive(), when rule_status_map[rule_id] holds a
real result, print the rule banner and the cached status and return it
without dispatching the test.
- Keep the reused result visible in the log - ideally tagged (e.g.
(cached)) so compliance reports still show which parent required the rule
and do not appear to lose coverage.
- Make it controllable, e.g. a
--no-rule-cache CLI option, so debugging a
single rule under a specific parent remains possible.
Impact if adopted
- Measured upper bound on this platform: 328 ms of a 555 ms suite (59%) is
repeat dispatch. Even a conservative idempotent-only policy covering
PCI_IN_02 and PCI_IN_16 recovers 257 ms (46%).
- On emulation this is significant: enabling these two rules alone pushed the
suite from 213.6 ms to 555.0 ms of simulated time and required raising
--time_limit from 400 ms to 800 ms. Each run occupies 72 emulator boards for
hours, so halving runtime materially increases regression throughput.
- Benefit is largest on emulation and silicon, where ECAM walks are slow; on
fast models the saving will be smaller.
Component:
val/src/rule_based_orchestrator.c(sysarch-acs)Type: Enhancement / runtime optimization
Severity: Low correctness risk, high runtime impact
Summary
The rule-based orchestrator re-executes a child rule once for every parent rule
that references it. Results are already recorded in
rule_status_map[], but themap is never consulted before execution, so identical work is repeated.
On our platform this roughly triples ACS wall time. In a full run that completed
at 555.0 ms of simulated time, 328.4 ms of the 506.1 ms rule span (65%) was
spent on second-and-later dispatches of rules that had already produced a
result.
Evidence
Worst offender is
PCI_IN_02("Check ECAM Memory accessibility"), which walksthe entire ECAM (buses 0x00-0xFF x 32 dev x 8 func):
PCI_IN_16("Check all 1's for out of range") behaves identically: fourdispatches of 31.2 ms each, 124.9 ms total, 93.7 ms of it redundant.
Together these two rules account for 341.5 ms of the 555.0 ms suite (62%).
Aggregate cost of repeat dispatches (2nd and later), same run:
Repetition is widespread, not limited to PCIe. Of 372 distinct rules observed:
285 dispatched once, 21 twice, 58 three times, and 8 four times.
Root cause
In
execute_rule_recursive():(
rule_reference_path_contains()) and explicit skips (is_rule_skipped()).There is no check for "this rule already produced a result in this run".
rule_status_map[]is therefore already populated with exactly the value acache would need. It is currently used only for reporting.
The result of a
BASE_RULEis a pure function of the rule and the PE count -the dispatch carries no parent-specific context:
This is what makes reuse sound in principle: a second execution under a
different parent is given identical inputs, so it can only produce an identical
result (absent side effects - see Risks).
Proposed enhancement
Consult
rule_status_map[]before executing a rule, and reuse a recordedresult instead of re-running the test.
Suggested shape:
rule_status_map[]atthe start of each run (it must not persist across runs).
execute_rule_recursive(), whenrule_status_map[rule_id]holds areal result, print the rule banner and the cached status and return it
without dispatching the test.
(cached)) so compliance reports still show which parent required the ruleand do not appear to lose coverage.
--no-rule-cacheCLI option, so debugging asingle rule under a specific parent remains possible.
Impact if adopted
repeat dispatch. Even a conservative idempotent-only policy covering
PCI_IN_02andPCI_IN_16recovers 257 ms (46%).suite from 213.6 ms to 555.0 ms of simulated time and required raising
--time_limitfrom 400 ms to 800 ms. Each run occupies 72 emulator boards forhours, so halving runtime materially increases regression throughput.
fast models the saving will be smaller.