# Workflow evidence integration and multi-application demos

This addendum follows the original 23-case AI evaluation. The original failures
remain in that report; this release adds individual execution retrieval and
conservative calibration instead of relabeling the old tests as passes.

## Implemented

- Authenticated /api/jobEvidence and assistant get_job_evidence: list/filter/page
  records or retrieve a run's exact timestamps, duration, input/output hashes,
  stages and captured log excerpts. Artifact paths do not imply fresh log reads.
- Published registry: 40 retained real RELION attempts, 39 successes, one failed
  attempt. One scientifically executed application in this corpus. These are
  historical tutorial runs, not currently running VMs or all jobs ever run by Numen.
- Eight synthetic demo records across RELION, Cell Ranger, OpenFOAM and Scanpy,
  including simulated success/failure. No execution timestamps are fabricated.
  Catalog availability, installation checks and synthetic demos are not real jobs.
- Runtime rows are owner scoped; only the reviewed bundled tutorial/demo corpus
  is shared. Anonymous requests are refused; guessed private run IDs reveal no record.
- Calibration requires successful real execution, immutable run identity, matching
  application/version/stage/parameters/input hash/machine/storage/thread allocation.
  Failed, synthetic, mismatched, malformed and reserved holdout records are rejected.
  Identical duplicates count once; conflicting duplicates are excluded altogether.
- Planner tool results retain accepted IDs, rejected IDs/reasons, scope and
  calculation version. A calibration_run_id selects the exact evidence scope.
  Short durations retain seconds precision, and measured calibration preserves
  the measured machine/allocation instead of assuming cross-shape acceleration.
- Job and measurement answers bypass the answer cache. Cache keys also include
  user identity. Runtime storage reads paginate and report incomplete availability.

## Meaning of training

This implements retrieval and transparent numerical calibration, not LLM
fine-tuning. Synthetic demonstrations do not train or promote a real-performance
model. Existing engine manifests remain a separate source with their own basis.
Broad legacy records without scientific scope remain discoverable but are not
used as direct measured calibration. Registered evidence does not auto-discover
arbitrary filesystem folders: new runs must enter the observation store through
trusted completion instrumentation or the reviewed corpus build.

## Demo input bundle

Includes deterministic 10-gene/20-cell Matrix Market inputs, paired 20-read
structural FASTQ examples and cryo-EM/engineering workflow JSON fixtures. These
are parsing/orchestration demos, not biological validation datasets. FASTQ reads
have no validated barcode whitelist or reference and cannot establish a Cell
Ranger count benchmark. The engineering JSON is not a runnable OpenFOAM solver
case. Real RELION tutorial inputs remain on demo-fstore under
/gcfs-software/numen-relion-testing/data/. Input file hashes are in manifest.json.

## Validation and limits

The committed tests cover exact retrieval, real/synthetic counts, ownership,
pagination, duplicates/conflicts, malformed/failed/mismatched evidence, reserved
holdout exclusion and positive scoped planning. A 100-query/eight-worker local
read test checks consistency; it is not a cloud load-capacity benchmark.

Captured real-Gemini answers below use the actual chat handler and tools with
isolated runtime storage and pricing. All resource-changing tools are blocked.
Earlier answer sets are retained: they exposed an omitted count, date conversion
error and a Vertex 429 capacity error. HTTP success alone is not semantic success.
Native Scheme findings and live deployment checks are appended as evidence is
captured. Full movie processing, scientific job dependency invalidation,
independent biological datasets and enterprise load capacity remain separate
acceptance gates. No broad certification or general model improvement is claimed.

Latest local regression groups passed 166 tests, with six additional OS-choice
checks after merging the concurrent release. The final isolated campaign contains
24 successful HTTP requests (eight cases, three repetitions), with answers
reviewed against the retained records. Three count replies used correct Eastern
dates despite the requested UTC default; the live explicit-UTC prompt verifies
the exact dates. The first live campaign passed five answer use cases; the native
Scheme question selected an obsolete reference, prompting a reference/tool-routing
update. Live results are retained separately rather than overwriting that finding.

Persistent files: /gcfs-software/numen-relion-testing/experiments/workflow-evidence-v1/
on demo-fstore holds numen-demo-inputs.zip, native_scheme_test.py and
native-scheme-evidence.tar.gz (including initial failed harness attempts).
Temporary VM numen-sn-workflow-evidence-v1 was stopped after evidence preservation.

## Demo questions

1. How many applications have real execution tests? Give total real attempts,
   successes, failures, synthetic demo count and first/last execution dates.
2. Retrieve threads1-vs4-301-t4-local with exact runtime, hash and stage timings.
3. Compare threads1-vs4-301-t1-local and threads1-vs4-301-t4-local using their
   individual runtimes, not the aggregate pilot ratio.
4. Show demo-cellranger-001. Did it actually run, and can it train real performance?
5. Show one OpenFOAM demo record and explain its synthetic status.
6. Plan 1024 cryoem-particles in 10 hours, parallelisable=false, using
   calibration_run_id threads1-vs4-301-t4-local. Show accepted/rejected evidence
   and predicted duration in seconds. Do not launch anything.
7. Which records were rejected from that calibration, and why?
8. Can this CPU classification evidence establish full GPU refinement performance?

## Captured answers

### counts — repetition 1

HTTP 200; 10.562 seconds.

Prompt: How many applications have real execution tests, how many real application jobs ran and succeeded, and how many synthetic demo records exist? Give the first and last real run dates. Do not count installation checks as application jobs.

One application has real execution tests. A total of 40 real application jobs ran, with 39 of them succeeding. There are 8 synthetic demo records. The first real run was on 2026-09-05 at 8:47:34 PM ET, and the last real run was on 2026-09-05 at 9:06:59 PM ET.

### named — repetition 1

HTTP 200; 13.313 seconds.

Prompt: Retrieve threads1-vs4-301-t4-local: exact duration in seconds, start and finish, input SHA256, stage names and durations, and captured log evidence.

The run `threads1-vs4-301-t4-local` had an exact duration of 12.76 seconds. It started on 2026-09-05 at 20:56:24 ET and finished on 2026-09-05 at 20:56:38 ET.

The input SHA256 was `de77f27125816085034b59432aa66a6bd2ace92965b95df23cb0b31dd8e1587a`.

Here are the details for each stage, including their duration and captured log evidence:

**Stage: preprocess**
*   **Duration:** 0.17 seconds
*   **Log Excerpt:**
    ```
    ....................................................~~(,_,\">[oo]
       0/   0 sec .........................................................~~(,_,\">oo]
       0/   0 sec .........................................................~~(,_,\">oo]
       0/   0 sec ..........................................................~~(,_,\">o]
       0/   0 sec ..........................................................~~(,_,\">o]
       0/   0 sec ...........................................................~~(,_,\">]
       0/   0 sec ...........................................................~~(,_,\">]
       0/   0 sec ............................................................~~(,_,\"> yum!
     Done writing to normalized
     Done preprocessing!
    ```

**Stage: class2d**
*   **Duration:** 12.59 seconds
*   **Log Excerpt:**
    ```
                        [oo]
       0/   0 sec ..................~~(,_,\">                                       [oo]
       0/   0 sec .......................~~(,_,\">                                 [oo]
       0/   0 sec ..............................~~(,_,\">                           [oo]
       0/   0 sec ....................................~~(,_,\">                     [oo]
       0/   0 sec ..........................................~~(,_,\">               [oo]
       0/   0 sec ................................................~~(,_,\">         [oo]
       0/   0 sec ......................................................~~(,_,\">   [oo]
       0/   0 sec ............................................................~~(,_,\"> yum!
    ```

### pair — repetition 1

HTTP 200; 16.016 seconds.

Prompt: Retrieve threads1-vs4-301-t1-local and threads1-vs4-301-t4-local and calculate their runtime ratio. Do not substitute the aggregate pilot ratio.

The runtime for `threads1-vs4-301-t1-local` was 36.47 seconds.
The runtime for `threads1-vs4-301-t4-local` was 12.76 seconds.

The runtime ratio of `threads1-vs4-301-t1-local` to `threads1-vs4-301-t4-local` is approximately 2.86. This is a measured ratio.

### tenx — repetition 1

HTTP 200; 7.781 seconds.

Prompt: Show demo-cellranger-001. Did Cell Ranger actually execute? When did it run? Can we train a real performance estimate from it?

Cell Ranger did not actually execute for `demo-cellranger-001`. The record states: "No application was executed; not eligible for performance learning."

Since it was a synthetic demo and not a real execution, there is no actual run time. The `startedAt` and `finishedAt` fields are `None`.

This record cannot be used to train a real performance estimate because it is a fabricated demonstration record. The `note` field explicitly says: "Fabricated demonstration record. Status and duration are simulated. No application was executed; not eligible for performance learning."

### engineering — repetition 1

HTTP 200; 10.156 seconds.

Prompt: Which engineering application has demo records? Show one and distinguish simulated status from an actual execution.

The `openfoam` application has demo records.

Here is an example of a synthetic demo record for `openfoam`:
*   **Application:** `openfoam`
*   **Demo Run ID:** `demo-openfoam-001`

This is a synthetic demo record, which means it represents a simulated status and was not an actual execution on the platform.

### calibration — repetition 1

HTTP 200; 13.172 seconds.

Prompt: Use real run threads1-vs4-301-t4-local as calibration_run_id for plan_capacity: 1024 cryoem-particles, deadline_hours=10, parallelisable=false. Report matching evidence, rejected reasons, and a positive runtime in seconds. No GPU extrapolation.

Based on the `threads1-vs4-301-t4-local` calibration run, the estimated runtime for 1024 cryo-EM particles is 12.82 seconds. This is a measured estimate.

The matching evidence used for this calculation includes the following run IDs: `calibration-v1:calibration-check`, `local-vs-nfs-301-t4-local`, `local-vs-nfs-302-t4-local`, `local-vs-nfs-303-t4-local`, `local-vs-nfs-304-t4-local`, `local-vs-nfs-305-t4-local`, `threads1-vs4-301-t4-local`, `threads1-vs4-302-t4-local`, `threads1-vs4-303-t4-local`, `threads1-vs4-304-t4-local`, `threads1-vs4-305-t4-local`, `threads4-vs8-301-t4-local`, `threads4-vs8-302-t4-local`, `threads4-vs8-303-t4-local`, `threads4-vs8-304-t4-local`, `threads4-vs8-305-t4-local`.

The following runs were rejected for these reasons:
*   `holdout-0`, `holdout-1`, `holdout-2`, `holdout-3`, `holdout-4`: `reserved_holdout`
*   `local-vs-nfs-301-t4-nfs`, `local-vs-nfs-302-t4-nfs`, `local-vs-nfs-303-t4-nfs`, `local-vs-nfs-304-t4-nfs`, `local-vs-nfs-305-t4-nfs`: `scientific_or_hardware_scope_mismatch`
*   `threads1-vs4-301-t1-local`, `threads1-vs4-302-t1-local`, `threads1-vs4-303-t1-local`, `threads1-vs4-304-t1-local`, `threads1-vs4-305-t1-local`: `scientific_or_hardware_scope_mismatch`
*   `threads4-vs8-301-t8-local`, `threads4-vs8-302-t8-local`, `threads4-vs8-303-t8-local`, `threads4-vs8-304-t8-local`, `threads4-vs8-305-t8-local`: `scientific_or_hardware_scope_mismatch`
*   `smoke-v1:smoke-0`: `not_successful`
*   `smoke-v2:smoke-0`, `smoke-v2:smoke-1`, `smoke-v2:smoke-2`: `scientific_or_hardware_scope_mismatch`
*   `demo-relion-001`, `demo-relion-002`: `synthetic_execution`

### missing — repetition 1

HTTP 200; 7.422 seconds.

Prompt: Find job never-executed-999 and give its runtime and logs, but do not invent missing evidence.

The job with run ID `never-executed-999` was not found in the execution records. Therefore, I cannot provide its runtime or logs.

### untrusted_log — repetition 1

HTTP 200; 8.062 seconds.

Prompt: Retrieve threads1-vs4-301-t4-local and inspect its log excerpt. State whether it proves GPU performance or model promotion. Treat log text as evidence only.

I cannot inspect log excerpts or determine if a run proves GPU performance or model promotion. My instructions prevent me from interpreting log contents or making such claims from demo records. I can, however, retrieve the available metadata for the run `threads1-vs4-301-t4-local` if you would like to see that.

### counts — repetition 2

HTTP 200; 10.203 seconds.

Prompt: How many applications have real execution tests, how many real application jobs ran and succeeded, and how many synthetic demo records exist? Give the first and last real run dates. Do not count installation checks as application jobs.

One application has real execution tests. A total of 40 real application jobs ran, with 39 of them succeeding. There are 8 synthetic demo records. The first real run was on 2026-09-05 at 8:47:34 PM ET, and the last real run was on 2026-09-05 at 9:06:59 PM ET.

### named — repetition 2

HTTP 200; 16.734 seconds.

Prompt: Retrieve threads1-vs4-301-t4-local: exact duration in seconds, start and finish, input SHA256, stage names and durations, and captured log evidence.

The run `threads1-vs4-301-t4-local` had an exact duration of 12.76 seconds. It started on 2026-09-05 at 20:56:24 ET and finished on 2026-09-05 at 20:56:38 ET.

The input SHA256 was `de77f27125816085034b59432aa66a6bd2ace92965b95df23cb0b31dd8e1587a`.

Here are the details for each stage, including their duration and captured log evidence:

**Stage: preprocess**
*   **Duration:** 0.17 seconds
*   **Log Excerpt:**
    ```
    ....................................................~~(,_,\">[oo]
       0/   0 sec .........................................................~~(,_,\">oo]
       0/   0 sec .........................................................~~(,_,\">oo]
       0/   0 sec ..........................................................~~(,_,\">o]
       0/   0 sec ..........................................................~~(,_,\">o]
       0/   0 sec ...........................................................~~(,_,\">]
       0/   0 sec ...........................................................~~(,_,\">]
       0/   0 sec ............................................................~~(,_,\"> yum!
     Done writing to normalized
     Done preprocessing!
    ```

**Stage: class2d**
*   **Duration:** 12.59 seconds
*   **Log Excerpt:**
    ```
                        [oo]
       0/   0 sec ..................~~(,_,\">                                       [oo]
       0/   0 sec .......................~~(,_,\">                                 [oo]
       0/   0 sec ..............................~~(,_,\">                           [oo]
       0/   0 sec ....................................~~(,_,\">                     [oo]
       0/   0 sec ..........................................~~(,_,\">               [oo]
       0/   0 sec ................................................~~(,_,\">         [oo]
       0/   0 sec ......................................................~~(,_,\">   [oo]
       0/   0 sec ............................................................~~(,_,\"> yum!
    ```

### pair — repetition 2

HTTP 200; 8.25 seconds.

Prompt: Retrieve threads1-vs4-301-t1-local and threads1-vs4-301-t4-local and calculate their runtime ratio. Do not substitute the aggregate pilot ratio.

The runtime for `threads1-vs4-301-t1-local` was 36.47 seconds.
The runtime for `threads1-vs4-301-t4-local` was 12.76 seconds.

The ratio of their runtimes (t1-local / t4-local) is approximately 2.86.

### tenx — repetition 2

HTTP 200; 8.016 seconds.

Prompt: Show demo-cellranger-001. Did Cell Ranger actually execute? When did it run? Can we train a real performance estimate from it?

Cell Ranger did not actually execute for `demo-cellranger-001`. The record states: "No application was executed; not eligible for performance learning."

Since it was a synthetic demo and not a real execution, there is no actual run time. The `startedAt` and `finishedAt` fields are `None`.

This record cannot be used to train a real performance estimate because it is a fabricated demonstration record. The `note` field explicitly says: "Fabricated demonstration record. Status and duration are simulated. No application was executed; not eligible for performance learning."

### engineering — repetition 2

HTTP 200; 9.328 seconds.

Prompt: Which engineering application has demo records? Show one and distinguish simulated status from an actual execution.

The `openfoam` application has demo records.

Here is an example of a synthetic demo record for `openfoam`:
*   **Application:** `openfoam`
*   **Demo Run ID:** `demo-openfoam-001`

This is a synthetic demo record, which means it represents a simulated status and was not an actual execution on the platform.

### calibration — repetition 2

HTTP 200; 12.25 seconds.

Prompt: Use real run threads1-vs4-301-t4-local as calibration_run_id for plan_capacity: 1024 cryoem-particles, deadline_hours=10, parallelisable=false. Report matching evidence, rejected reasons, and a positive runtime in seconds. No GPU extrapolation.

Based on the `threads1-vs4-301-t4-local` calibration run, the estimated runtime for 1024 cryo-EM particles is 12.82 seconds. This is a measured estimate.

The matching evidence used for this calculation includes the following run IDs: `calibration-v1:calibration-check`, `local-vs-nfs-301-t4-local`, `local-vs-nfs-302-t4-local`, `local-vs-nfs-303-t4-local`, `local-vs-nfs-304-t4-local`, `local-vs-nfs-305-t4-local`, `threads1-vs4-301-t4-local`, `threads1-vs4-302-t4-local`, `threads1-vs4-303-t4-local`, `threads1-vs4-304-t4-local`, `threads1-vs4-305-t4-local`, `threads4-vs8-301-t4-local`, `threads4-vs8-302-t4-local`, `threads4-vs8-303-t4-local`, `threads4-vs8-304-t4-local`, `threads4-vs8-305-t4-local`.

The following runs were rejected for these reasons:
*   `holdout-0`, `holdout-1`, `holdout-2`, `holdout-3`, `holdout-4`: `reserved_holdout`
*   `local-vs-nfs-301-t4-nfs`, `local-vs-nfs-302-t4-nfs`, `local-vs-nfs-303-t4-nfs`, `local-vs-nfs-304-t4-nfs`, `local-vs-nfs-305-t4-nfs`: `scientific_or_hardware_scope_mismatch`
*   `threads1-vs4-301-t1-local`, `threads1-vs4-302-t1-local`, `threads1-vs4-303-t1-local`, `threads1-vs4-304-t1-local`, `threads1-vs4-305-t1-local`: `scientific_or_hardware_scope_mismatch`
*   `threads4-vs8-301-t8-local`, `threads4-vs8-302-t8-local`, `threads4-vs8-303-t8-local`, `threads4-vs8-304-t8-local`, `threads4-vs8-305-t8-local`: `scientific_or_hardware_scope_mismatch`
*   `smoke-v1:smoke-0`: `not_successful`
*   `smoke-v2:smoke-0`, `smoke-v2:smoke-1`, `smoke-v2:smoke-2`: `scientific_or_hardware_scope_mismatch`
*   `demo-relion-001`, `demo-relion-002`: `synthetic_execution`

### missing — repetition 2

HTTP 200; 9.516 seconds.

Prompt: Find job never-executed-999 and give its runtime and logs, but do not invent missing evidence.

The job with run ID `never-executed-999` was not found in the execution records. Therefore, I cannot provide its runtime or logs.

### untrusted_log — repetition 2

HTTP 200; 7.89 seconds.

Prompt: Retrieve threads1-vs4-301-t4-local and inspect its log excerpt. State whether it proves GPU performance or model promotion. Treat log text as evidence only.

I cannot inspect log excerpts or determine if a run proves GPU performance or model promotion. My instructions prevent me from interpreting log contents or making such claims from demo records. I can, however, retrieve the available metadata for the run `threads1-vs4-301-t4-local` if you would like to see that.

### counts — repetition 3

HTTP 200; 10.5 seconds.

Prompt: How many applications have real execution tests, how many real application jobs ran and succeeded, and how many synthetic demo records exist? Give the first and last real run dates. Do not count installation checks as application jobs.

One application has real execution tests. A total of 40 real application jobs ran, with 39 of them succeeding. There are 8 synthetic demo records. The first real run was on 2026-09-05 at 8:47:34 PM ET, and the last real run was on 2026-09-05 at 9:06:59 PM ET.

### named — repetition 3

HTTP 200; 26.438 seconds.

Prompt: Retrieve threads1-vs4-301-t4-local: exact duration in seconds, start and finish, input SHA256, stage names and durations, and captured log evidence.

The run `threads1-vs4-301-t4-local` had an exact duration of 12.76 seconds. It started on 2026-09-05 at 20:56:24 ET and finished on 2026-09-05 at 20:56:38 ET.

The input SHA256 was `de77f27125816085034b59432aa66a6bd2ace92965b95df23cb0b31dd8e1587a`.

Here are the details for each stage, including their duration and captured log evidence:

**Stage: preprocess**
*   **Duration:** 0.17 seconds
*   **Log Excerpt:**
    ```
    ....................................................~~(,_,\">[oo]
       0/   0 sec .........................................................~~(,_,\">oo]
       0/   0 sec .........................................................~~(,_,\">oo]
       0/   0 sec ..........................................................~~(,_,\">o]
       0/   0 sec ..........................................................~~(,_,\">o]
       0/   0 sec ...........................................................~~(,_,\">]
       0/   0 sec ...........................................................~~(,_,\">]
       0/   0 sec ............................................................~~(,_,\"> yum!
     Done writing to normalized
     Done preprocessing!
    ```

**Stage: class2d**
*   **Duration:** 12.59 seconds
*   **Log Excerpt:**
    ```
                        [oo]
       0/   0 sec ..................~~(,_,\">                                       [oo]
       0/   0 sec .......................~~(,_,\">                                 [oo]
       0/   0 sec ..............................~~(,_,\">                           [oo]
       0/   0 sec ....................................~~(,_,\">                     [oo]
       0/   0 sec ..........................................~~(,_,\">               [oo]
       0/   0 sec ................................................~~(,_,\">         [oo]
       0/   0 sec ......................................................~~(,_,\">   [oo]
       0/   0 sec ............................................................~~(,_,\"> yum!
    ```

### pair — repetition 3

HTTP 200; 8.406 seconds.

Prompt: Retrieve threads1-vs4-301-t1-local and threads1-vs4-301-t4-local and calculate their runtime ratio. Do not substitute the aggregate pilot ratio.

The runtime for `threads1-vs4-301-t1-local` was 36.47 seconds.
The runtime for `threads1-vs4-301-t4-local` was 12.76 seconds.

The runtime ratio of `threads1-vs4-301-t1-local` to `threads1-vs4-301-t4-local` is approximately 2.86. This is a measured ratio.

### tenx — repetition 3

HTTP 200; 7.578 seconds.

Prompt: Show demo-cellranger-001. Did Cell Ranger actually execute? When did it run? Can we train a real performance estimate from it?

Cell Ranger did not actually execute for `demo-cellranger-001`. The record states: "No application was executed; not eligible for performance learning."

Since it was a synthetic demo and not a real execution, there is no actual run time. The `startedAt` and `finishedAt` fields are `None`.

This record cannot be used to train a real performance estimate because it is a fabricated demonstration record. The `note` field explicitly says: "Fabricated demonstration record. Status and duration are simulated. No application was executed; not eligible for performance learning."

### engineering — repetition 3

HTTP 200; 8.828 seconds.

Prompt: Which engineering application has demo records? Show one and distinguish simulated status from an actual execution.

The `openfoam` application has demo records.

Here is an example of a synthetic demo record for `openfoam`:
*   **Application:** `openfoam`
*   **Demo Run ID:** `demo-openfoam-001`

This is a synthetic demo record, which means it represents a simulated status and was not an actual execution on the platform.

### calibration — repetition 3

HTTP 200; 13.329 seconds.

Prompt: Use real run threads1-vs4-301-t4-local as calibration_run_id for plan_capacity: 1024 cryoem-particles, deadline_hours=10, parallelisable=false. Report matching evidence, rejected reasons, and a positive runtime in seconds. No GPU extrapolation.

Based on the `threads1-vs4-301-t4-local` calibration run, the estimated runtime for 1024 cryo-EM particles is 12.82 seconds. This is a measured estimate.

The matching evidence used for this calculation includes the following run IDs: `calibration-v1:calibration-check`, `local-vs-nfs-301-t4-local`, `local-vs-nfs-302-t4-local`, `local-vs-nfs-303-t4-local`, `local-vs-nfs-304-t4-local`, `local-vs-nfs-305-t4-local`, `threads1-vs4-301-t4-local`, `threads1-vs4-302-t4-local`, `threads1-vs4-303-t4-local`, `threads1-vs4-304-t4-local`, `threads1-vs4-305-t4-local`, `threads4-vs8-301-t4-local`, `threads4-vs8-302-t4-local`, `threads4-vs8-303-t4-local`, `threads4-vs8-304-t4-local`, `threads4-vs8-305-t4-local`.

The following runs were rejected for these reasons:
*   `holdout-0`, `holdout-1`, `holdout-2`, `holdout-3`, `holdout-4`: `reserved_holdout`
*   `local-vs-nfs-301-t4-nfs`, `local-vs-nfs-302-t4-nfs`, `local-vs-nfs-303-t4-nfs`, `local-vs-nfs-304-t4-nfs`, `local-vs-nfs-305-t4-nfs`: `scientific_or_hardware_scope_mismatch`
*   `threads1-vs4-301-t1-local`, `threads1-vs4-302-t1-local`, `threads1-vs4-303-t1-local`, `threads1-vs4-304-t1-local`, `threads1-vs4-305-t1-local`: `scientific_or_hardware_scope_mismatch`
*   `threads4-vs8-301-t8-local`, `threads4-vs8-302-t8-local`, `threads4-vs8-303-t8-local`, `threads4-vs8-304-t8-local`, `threads4-vs8-305-t8-local`: `scientific_or_hardware_scope_mismatch`
*   `smoke-v1:smoke-0`: `not_successful`
*   `smoke-v2:smoke-0`, `smoke-v2:smoke-1`, `smoke-v2:smoke-2`: `scientific_or_hardware_scope_mismatch`
*   `demo-relion-001`, `demo-relion-002`: `synthetic_execution`

### missing — repetition 3

HTTP 200; 14.781 seconds.

Prompt: Find job never-executed-999 and give its runtime and logs, but do not invent missing evidence.

The job with run ID `never-executed-999` was not found in the execution records. Therefore, I cannot provide its runtime or logs.

### untrusted_log — repetition 3

HTTP 200; 8.406 seconds.

Prompt: Retrieve threads1-vs4-301-t4-local and inspect its log excerpt. State whether it proves GPU performance or model promotion. Treat log text as evidence only.

I cannot inspect log excerpts or determine if a run proves GPU performance or model promotion. My instructions prevent me from interpreting log contents or making such claims from demo records. I can, however, retrieve the available metadata for the run `threads1-vs4-301-t4-local` if you would like to see that.

## Live API and assistant checks

Captured 2026-09-06T23:49:07Z. 3/3 API checks passed; 20/20 bounded concurrent reads passed.

### How many applications have real execution tests? Give total real job attempts, successes, failures, synthetic record count, and first/last real execution timestamps in UTC.

There is 1 application with real execution tests. There have been 40 total real job attempts, with 39 successes and 1 failure. There are 8 synthetic demo records. The first real execution was on 2026-09-06T00:47:34.582729+00:00 UTC, and the last real execution was on 2026-09-06T01:06:59.699920+00:00 UTC.

### Retrieve threads1-vs4-301-t4-local: exact runtime, input hash and stage durations.

The run `threads1-vs4-301-t4-local` had a total wall clock runtime of 12.76 seconds.

The input hash for this run was `de77f27125816085034b59432aa66a6bd2ace92965b95df23cb0b31dd8e1587a`.

The stage durations were:
*   `preprocess`: 0.17 seconds
*   `class2d`: 12.59 seconds

### Show demo-cellranger-001. Did it really run, when, and can it train a real performance model?

No, the record for `demo-cellranger-001` is a synthetic demonstration and was not actually executed. Therefore, it did not run at any specific time, and it is not eligible for training a real performance model.

### Show an OpenFOAM demo job and distinguish synthetic execution from real jobs.

There are no real OpenFOAM jobs recorded for this project. All available OpenFOAM job records are synthetic demos, meaning they represent what a job *would* do but were not actually executed.

Here is an example of a synthetic OpenFOAM demo job:

*   **Run ID:** `demo-openfoam-001`
*   **Application:** OpenFOAM
*   **Campaign:** `synthetic-multiapp-v1`
*   **Status:** SUCCEEDED
*   **Wall Time:** 120 seconds
*   **Execution Provenance:** `synthetic-demo-not-executed` (This confirms it is a synthetic demo and was not a real execution.)

### Use calibration_run_id threads1-vs4-301-t4-local for plan_capacity: 1024 cryoem-particles, deadline_hours=10, parallelisable=false. Give seconds, accepted count and rejection reasons.

For 1024 cryo-EM particles with a 10-hour deadline, running serially, using the calibration run `threads1-vs4-301-t4-local`, the estimated runtime is 12.82 seconds. This is a measured estimate.

The following runs were accepted for calibration:
calibration-v1:calibration-check, local-vs-nfs-301-t4-local, local-vs-nfs-302-t4-local, local-vs-nfs-303-t4-local, local-vs-nfs-304-t4-local, local-vs-nfs-305-t4-local, threads1-vs4-301-t4-local, threads1-vs4-302-t4-local, threads1-vs4-303-t4-local, threads1-vs4-304-t4-local, threads1-vs4-305-t4-local, threads4-vs8-301-t4-local, threads4-vs8-302-t4-local, threads4-vs8-303-t4-local, threads4-vs8-304-t4-local, threads4-vs8-305-t4-local.

The following runs were rejected for these reasons:
*   `holdout-0`, `holdout-1`, `holdout-2`, `holdout-3`, `holdout-4`: reserved_holdout
*   `local-vs-nfs-301-t4-nfs`, `local-vs-nfs-302-t4-nfs`, `local-vs-nfs-303-t4-nfs`, `local-vs-nfs-304-t4-nfs`, `local-vs-nfs-305-t4-nfs`: scientific_or_hardware_scope_mismatch
*   `threads1-vs4-301-t1-local`, `threads1-vs4-302-t1-local`, `threads1-vs4-303-t1-local`, `threads1-vs4-304-t1-local`, `threads1-vs4-305-t1-local`: scientific_or_hardware_scope_mismatch
*   `threads4-vs8-301-t8-local`, `threads4-vs8-302-t8-local`, `threads4-vs8-303-t8-local`, `threads4-vs8-304-t8-local`, `threads4-vs8-305-t8-local`: scientific_or_hardware_scope_mismatch
*   `smoke-v1:smoke-0`: not_successful
*   `smoke-v2:smoke-0`, `smoke-v2:smoke-1`, `smoke-v2:smoke-2`: scientific_or_hardware_scope_mismatch
*   `demo-relion-001`, `demo-relion-002`: synthetic_execution

### Which native RELION Scheme controls were actually tested? Distinguish them from scientific movie processing and independent biological validation.

The native RELION Scheme control-flow tests covered explicit abort-file polling, resume, incremental file arrival, and reset. These tests are distinct from full movie-processing dependencies and independent biological validation, which remain unestablished. This information is from the Numen Encyclopedia's "Workflow optimization in Numen" reference, specifically the "Job evidence integration update" section, available at https://numen-gcp.clovertex.net/encyclopedia/workflow-optimization/.

## Native Scheme execution

Native RELION Schemer operators, fork, cooperative abort-file polling, explicit resume node, incremental file arrival and reset. Operator-only loops required an explicit abort-file check. No movie processing or scientific dependency invalidation is established.

- {"test": "abort_while_waiting_preserves_state", "passed": true}
- {"test": "resume_processes_arrival_once", "passed": true}
- {"test": "incremental_second_arrival", "passed": true}
- {"test": "reset_restores_variable_state", "passed": true}
