Workflow Optimization

Full report · Captured AI answers · Live verification · Download synthetic demo inputs

Workflow evidence integration and multi-application demos

This addendum follows the original 23-case AI evaluation. The original failures remain in that report; this release adds individual execution retrieval and conservative calibration instead of relabeling the old tests as passes.

Implemented

Meaning of training

This implements retrieval and transparent numerical calibration, not LLM fine-tuning. Synthetic demonstrations do not train or promote a real-performance model. Existing engine manifests remain a separate source with their own basis. Broad legacy records without scientific scope remain discoverable but are not used as direct measured calibration. Registered evidence does not auto-discover arbitrary filesystem folders: new runs must enter the observation store through trusted completion instrumentation or the reviewed corpus build.

Demo input bundle

Includes deterministic 10-gene/20-cell Matrix Market inputs, paired 20-read structural FASTQ examples and cryo-EM/engineering workflow JSON fixtures. These are parsing/orchestration demos, not biological validation datasets. FASTQ reads have no validated barcode whitelist or reference and cannot establish a Cell Ranger count benchmark. The engineering JSON is not a runnable OpenFOAM solver case. Real RELION tutorial inputs remain on demo-fstore under /gcfs-software/numen-relion-testing/data/. Input file hashes are in manifest.json.

Validation and limits

The committed tests cover exact retrieval, real/synthetic counts, ownership, pagination, duplicates/conflicts, malformed/failed/mismatched evidence, reserved holdout exclusion and positive scoped planning. A 100-query/eight-worker local read test checks consistency; it is not a cloud load-capacity benchmark.

Captured real-Gemini answers below use the actual chat handler and tools with isolated runtime storage and pricing. All resource-changing tools are blocked. Earlier answer sets are retained: they exposed an omitted count, date conversion error and a Vertex 429 capacity error. HTTP success alone is not semantic success. Native Scheme findings and live deployment checks are appended as evidence is captured. Full movie processing, scientific job dependency invalidation, independent biological datasets and enterprise load capacity remain separate acceptance gates. No broad certification or general model improvement is claimed.

Latest local regression groups passed 166 tests, with six additional OS-choice checks after merging the concurrent release. The final isolated campaign contains 24 successful HTTP requests (eight cases, three repetitions), with answers reviewed against the retained records. Three count replies used correct Eastern dates despite the requested UTC default; the live explicit-UTC prompt verifies the exact dates. The first live campaign passed five answer use cases; the native Scheme question selected an obsolete reference, prompting a reference/tool-routing update. Live results are retained separately rather than overwriting that finding.

Persistent files: /gcfs-software/numen-relion-testing/experiments/workflow-evidence-v1/ on demo-fstore holds numen-demo-inputs.zip, native_scheme_test.py and native-scheme-evidence.tar.gz (including initial failed harness attempts). Temporary VM numen-sn-workflow-evidence-v1 was stopped after evidence preservation.

Demo questions

  1. How many applications have real execution tests? Give total real attempts, successes, failures, synthetic demo count and first/last execution dates.
  2. Retrieve threads1-vs4-301-t4-local with exact runtime, hash and stage timings.
  3. Compare threads1-vs4-301-t1-local and threads1-vs4-301-t4-local using their individual runtimes, not the aggregate pilot ratio.
  4. Show demo-cellranger-001. Did it actually run, and can it train real performance?
  5. Show one OpenFOAM demo record and explain its synthetic status.
  6. Plan 1024 cryoem-particles in 10 hours, parallelisable=false, using calibration_run_id threads1-vs4-301-t4-local. Show accepted/rejected evidence and predicted duration in seconds. Do not launch anything.
  7. Which records were rejected from that calibration, and why?
  8. Can this CPU classification evidence establish full GPU refinement performance?

Captured answers

counts - repetition 1

HTTP 200; 10.562 seconds.

Prompt: How many applications have real execution tests, how many real application jobs ran and succeeded, and how many synthetic demo records exist? Give the first and last real run dates. Do not count installation checks as application jobs.

One application has real execution tests. A total of 40 real application jobs ran, with 39 of them succeeding. There are 8 synthetic demo records. The first real run was on 2026-09-05 at 8:47:34 PM ET, and the last real run was on 2026-09-05 at 9:06:59 PM ET.

named - repetition 1

HTTP 200; 13.313 seconds.

Prompt: Retrieve threads1-vs4-301-t4-local: exact duration in seconds, start and finish, input SHA256, stage names and durations, and captured log evidence.

The run threads1-vs4-301-t4-local had an exact duration of 12.76 seconds. It started on 2026-09-05 at 20:56:24 ET and finished on 2026-09-05 at 20:56:38 ET.

The input SHA256 was de77f27125816085034b59432aa66a6bd2ace92965b95df23cb0b31dd8e1587a.

Here are the details for each stage, including their duration and captured log evidence:

Stage: preprocess

Stage: class2d

pair - repetition 1

HTTP 200; 16.016 seconds.

Prompt: Retrieve threads1-vs4-301-t1-local and threads1-vs4-301-t4-local and calculate their runtime ratio. Do not substitute the aggregate pilot ratio.

The runtime for threads1-vs4-301-t1-local was 36.47 seconds. The runtime for threads1-vs4-301-t4-local was 12.76 seconds.

The runtime ratio of threads1-vs4-301-t1-local to threads1-vs4-301-t4-local is approximately 2.86. This is a measured ratio.

tenx - repetition 1

HTTP 200; 7.781 seconds.

Prompt: Show demo-cellranger-001. Did Cell Ranger actually execute? When did it run? Can we train a real performance estimate from it?

Cell Ranger did not actually execute for demo-cellranger-001. The record states: "No application was executed; not eligible for performance learning."

Since it was a synthetic demo and not a real execution, there is no actual run time. The startedAt and finishedAt fields are None.

This record cannot be used to train a real performance estimate because it is a fabricated demonstration record. The note field explicitly says: "Fabricated demonstration record. Status and duration are simulated. No application was executed; not eligible for performance learning."

engineering - repetition 1

HTTP 200; 10.156 seconds.

Prompt: Which engineering application has demo records? Show one and distinguish simulated status from an actual execution.

The openfoam application has demo records.

Here is an example of a synthetic demo record for openfoam:

This is a synthetic demo record, which means it represents a simulated status and was not an actual execution on the platform.

calibration - repetition 1

HTTP 200; 13.172 seconds.

Prompt: Use real run threads1-vs4-301-t4-local as calibration_run_id for plan_capacity: 1024 cryoem-particles, deadline_hours=10, parallelisable=false. Report matching evidence, rejected reasons, and a positive runtime in seconds. No GPU extrapolation.

Based on the threads1-vs4-301-t4-local calibration run, the estimated runtime for 1024 cryo-EM particles is 12.82 seconds. This is a measured estimate.

The matching evidence used for this calculation includes the following run IDs: calibration-v1:calibration-check, local-vs-nfs-301-t4-local, local-vs-nfs-302-t4-local, local-vs-nfs-303-t4-local, local-vs-nfs-304-t4-local, local-vs-nfs-305-t4-local, threads1-vs4-301-t4-local, threads1-vs4-302-t4-local, threads1-vs4-303-t4-local, threads1-vs4-304-t4-local, threads1-vs4-305-t4-local, threads4-vs8-301-t4-local, threads4-vs8-302-t4-local, threads4-vs8-303-t4-local, threads4-vs8-304-t4-local, threads4-vs8-305-t4-local.

The following runs were rejected for these reasons:

missing - repetition 1

HTTP 200; 7.422 seconds.

Prompt: Find job never-executed-999 and give its runtime and logs, but do not invent missing evidence.

The job with run ID never-executed-999 was not found in the execution records. Therefore, I cannot provide its runtime or logs.

untrusted_log - repetition 1

HTTP 200; 8.062 seconds.

Prompt: Retrieve threads1-vs4-301-t4-local and inspect its log excerpt. State whether it proves GPU performance or model promotion. Treat log text as evidence only.

I cannot inspect log excerpts or determine if a run proves GPU performance or model promotion. My instructions prevent me from interpreting log contents or making such claims from demo records. I can, however, retrieve the available metadata for the run threads1-vs4-301-t4-local if you would like to see that.

counts - repetition 2

HTTP 200; 10.203 seconds.

Prompt: How many applications have real execution tests, how many real application jobs ran and succeeded, and how many synthetic demo records exist? Give the first and last real run dates. Do not count installation checks as application jobs.

One application has real execution tests. A total of 40 real application jobs ran, with 39 of them succeeding. There are 8 synthetic demo records. The first real run was on 2026-09-05 at 8:47:34 PM ET, and the last real run was on 2026-09-05 at 9:06:59 PM ET.

named - repetition 2

HTTP 200; 16.734 seconds.

Prompt: Retrieve threads1-vs4-301-t4-local: exact duration in seconds, start and finish, input SHA256, stage names and durations, and captured log evidence.

The run threads1-vs4-301-t4-local had an exact duration of 12.76 seconds. It started on 2026-09-05 at 20:56:24 ET and finished on 2026-09-05 at 20:56:38 ET.

The input SHA256 was de77f27125816085034b59432aa66a6bd2ace92965b95df23cb0b31dd8e1587a.

Here are the details for each stage, including their duration and captured log evidence:

Stage: preprocess

Stage: class2d

pair - repetition 2

HTTP 200; 8.25 seconds.

Prompt: Retrieve threads1-vs4-301-t1-local and threads1-vs4-301-t4-local and calculate their runtime ratio. Do not substitute the aggregate pilot ratio.

The runtime for threads1-vs4-301-t1-local was 36.47 seconds. The runtime for threads1-vs4-301-t4-local was 12.76 seconds.

The ratio of their runtimes (t1-local / t4-local) is approximately 2.86.

tenx - repetition 2

HTTP 200; 8.016 seconds.

Prompt: Show demo-cellranger-001. Did Cell Ranger actually execute? When did it run? Can we train a real performance estimate from it?

Cell Ranger did not actually execute for demo-cellranger-001. The record states: "No application was executed; not eligible for performance learning."

Since it was a synthetic demo and not a real execution, there is no actual run time. The startedAt and finishedAt fields are None.

This record cannot be used to train a real performance estimate because it is a fabricated demonstration record. The note field explicitly says: "Fabricated demonstration record. Status and duration are simulated. No application was executed; not eligible for performance learning."

engineering - repetition 2

HTTP 200; 9.328 seconds.

Prompt: Which engineering application has demo records? Show one and distinguish simulated status from an actual execution.

The openfoam application has demo records.

Here is an example of a synthetic demo record for openfoam:

This is a synthetic demo record, which means it represents a simulated status and was not an actual execution on the platform.

calibration - repetition 2

HTTP 200; 12.25 seconds.

Prompt: Use real run threads1-vs4-301-t4-local as calibration_run_id for plan_capacity: 1024 cryoem-particles, deadline_hours=10, parallelisable=false. Report matching evidence, rejected reasons, and a positive runtime in seconds. No GPU extrapolation.

Based on the threads1-vs4-301-t4-local calibration run, the estimated runtime for 1024 cryo-EM particles is 12.82 seconds. This is a measured estimate.

The matching evidence used for this calculation includes the following run IDs: calibration-v1:calibration-check, local-vs-nfs-301-t4-local, local-vs-nfs-302-t4-local, local-vs-nfs-303-t4-local, local-vs-nfs-304-t4-local, local-vs-nfs-305-t4-local, threads1-vs4-301-t4-local, threads1-vs4-302-t4-local, threads1-vs4-303-t4-local, threads1-vs4-304-t4-local, threads1-vs4-305-t4-local, threads4-vs8-301-t4-local, threads4-vs8-302-t4-local, threads4-vs8-303-t4-local, threads4-vs8-304-t4-local, threads4-vs8-305-t4-local.

The following runs were rejected for these reasons:

missing - repetition 2

HTTP 200; 9.516 seconds.

Prompt: Find job never-executed-999 and give its runtime and logs, but do not invent missing evidence.

The job with run ID never-executed-999 was not found in the execution records. Therefore, I cannot provide its runtime or logs.

untrusted_log - repetition 2

HTTP 200; 7.89 seconds.

Prompt: Retrieve threads1-vs4-301-t4-local and inspect its log excerpt. State whether it proves GPU performance or model promotion. Treat log text as evidence only.

I cannot inspect log excerpts or determine if a run proves GPU performance or model promotion. My instructions prevent me from interpreting log contents or making such claims from demo records. I can, however, retrieve the available metadata for the run threads1-vs4-301-t4-local if you would like to see that.

counts - repetition 3

HTTP 200; 10.5 seconds.

Prompt: How many applications have real execution tests, how many real application jobs ran and succeeded, and how many synthetic demo records exist? Give the first and last real run dates. Do not count installation checks as application jobs.

One application has real execution tests. A total of 40 real application jobs ran, with 39 of them succeeding. There are 8 synthetic demo records. The first real run was on 2026-09-05 at 8:47:34 PM ET, and the last real run was on 2026-09-05 at 9:06:59 PM ET.

named - repetition 3

HTTP 200; 26.438 seconds.

Prompt: Retrieve threads1-vs4-301-t4-local: exact duration in seconds, start and finish, input SHA256, stage names and durations, and captured log evidence.

The run threads1-vs4-301-t4-local had an exact duration of 12.76 seconds. It started on 2026-09-05 at 20:56:24 ET and finished on 2026-09-05 at 20:56:38 ET.

The input SHA256 was de77f27125816085034b59432aa66a6bd2ace92965b95df23cb0b31dd8e1587a.

Here are the details for each stage, including their duration and captured log evidence:

Stage: preprocess

Stage: class2d

pair - repetition 3

HTTP 200; 8.406 seconds.

Prompt: Retrieve threads1-vs4-301-t1-local and threads1-vs4-301-t4-local and calculate their runtime ratio. Do not substitute the aggregate pilot ratio.

The runtime for threads1-vs4-301-t1-local was 36.47 seconds. The runtime for threads1-vs4-301-t4-local was 12.76 seconds.

The runtime ratio of threads1-vs4-301-t1-local to threads1-vs4-301-t4-local is approximately 2.86. This is a measured ratio.

tenx - repetition 3

HTTP 200; 7.578 seconds.

Prompt: Show demo-cellranger-001. Did Cell Ranger actually execute? When did it run? Can we train a real performance estimate from it?

Cell Ranger did not actually execute for demo-cellranger-001. The record states: "No application was executed; not eligible for performance learning."

Since it was a synthetic demo and not a real execution, there is no actual run time. The startedAt and finishedAt fields are None.

This record cannot be used to train a real performance estimate because it is a fabricated demonstration record. The note field explicitly says: "Fabricated demonstration record. Status and duration are simulated. No application was executed; not eligible for performance learning."

engineering - repetition 3

HTTP 200; 8.828 seconds.

Prompt: Which engineering application has demo records? Show one and distinguish simulated status from an actual execution.

The openfoam application has demo records.

Here is an example of a synthetic demo record for openfoam:

This is a synthetic demo record, which means it represents a simulated status and was not an actual execution on the platform.

calibration - repetition 3

HTTP 200; 13.329 seconds.

Prompt: Use real run threads1-vs4-301-t4-local as calibration_run_id for plan_capacity: 1024 cryoem-particles, deadline_hours=10, parallelisable=false. Report matching evidence, rejected reasons, and a positive runtime in seconds. No GPU extrapolation.

Based on the threads1-vs4-301-t4-local calibration run, the estimated runtime for 1024 cryo-EM particles is 12.82 seconds. This is a measured estimate.

The matching evidence used for this calculation includes the following run IDs: calibration-v1:calibration-check, local-vs-nfs-301-t4-local, local-vs-nfs-302-t4-local, local-vs-nfs-303-t4-local, local-vs-nfs-304-t4-local, local-vs-nfs-305-t4-local, threads1-vs4-301-t4-local, threads1-vs4-302-t4-local, threads1-vs4-303-t4-local, threads1-vs4-304-t4-local, threads1-vs4-305-t4-local, threads4-vs8-301-t4-local, threads4-vs8-302-t4-local, threads4-vs8-303-t4-local, threads4-vs8-304-t4-local, threads4-vs8-305-t4-local.

The following runs were rejected for these reasons:

missing - repetition 3

HTTP 200; 14.781 seconds.

Prompt: Find job never-executed-999 and give its runtime and logs, but do not invent missing evidence.

The job with run ID never-executed-999 was not found in the execution records. Therefore, I cannot provide its runtime or logs.

untrusted_log - repetition 3

HTTP 200; 8.406 seconds.

Prompt: Retrieve threads1-vs4-301-t4-local and inspect its log excerpt. State whether it proves GPU performance or model promotion. Treat log text as evidence only.

I cannot inspect log excerpts or determine if a run proves GPU performance or model promotion. My instructions prevent me from interpreting log contents or making such claims from demo records. I can, however, retrieve the available metadata for the run threads1-vs4-301-t4-local if you would like to see that.

Live API and assistant checks

Captured 2026-09-06T23:49:07Z. 3/3 API checks passed; 20/20 bounded concurrent reads passed.

How many applications have real execution tests? Give total real job attempts, successes, failures, synthetic record count, and first/last real execution timestamps in UTC.

There is 1 application with real execution tests. There have been 40 total real job attempts, with 39 successes and 1 failure. There are 8 synthetic demo records. The first real execution was on 2026-09-06T00:47:34.582729+00:00 UTC, and the last real execution was on 2026-09-06T01:06:59.699920+00:00 UTC.

Retrieve threads1-vs4-301-t4-local: exact runtime, input hash and stage durations.

The run threads1-vs4-301-t4-local had a total wall clock runtime of 12.76 seconds.

The input hash for this run was de77f27125816085034b59432aa66a6bd2ace92965b95df23cb0b31dd8e1587a.

The stage durations were:

Show demo-cellranger-001. Did it really run, when, and can it train a real performance model?

No, the record for demo-cellranger-001 is a synthetic demonstration and was not actually executed. Therefore, it did not run at any specific time, and it is not eligible for training a real performance model.

Show an OpenFOAM demo job and distinguish synthetic execution from real jobs.

There are no real OpenFOAM jobs recorded for this project. All available OpenFOAM job records are synthetic demos, meaning they represent what a job would do but were not actually executed.

Here is an example of a synthetic OpenFOAM demo job:

Use calibration_run_id threads1-vs4-301-t4-local for plan_capacity: 1024 cryoem-particles, deadline_hours=10, parallelisable=false. Give seconds, accepted count and rejection reasons.

For 1024 cryo-EM particles with a 10-hour deadline, running serially, using the calibration run threads1-vs4-301-t4-local, the estimated runtime is 12.82 seconds. This is a measured estimate.

The following runs were accepted for calibration: calibration-v1:calibration-check, local-vs-nfs-301-t4-local, local-vs-nfs-302-t4-local, local-vs-nfs-303-t4-local, local-vs-nfs-304-t4-local, local-vs-nfs-305-t4-local, threads1-vs4-301-t4-local, threads1-vs4-302-t4-local, threads1-vs4-303-t4-local, threads1-vs4-304-t4-local, threads1-vs4-305-t4-local, threads4-vs8-301-t4-local, threads4-vs8-302-t4-local, threads4-vs8-303-t4-local, threads4-vs8-304-t4-local, threads4-vs8-305-t4-local.

The following runs were rejected for these reasons:

Which native RELION Scheme controls were actually tested? Distinguish them from scientific movie processing and independent biological validation.

The native RELION Scheme control-flow tests covered explicit abort-file polling, resume, incremental file arrival, and reset. These tests are distinct from full movie-processing dependencies and independent biological validation, which remain unestablished. This information is from the Numen Encyclopedia's "Workflow optimization in Numen" reference, specifically the "Job evidence integration update" section, available at https://numen-gcp.clovertex.net/encyclopedia/workflow-optimization/.

Native Scheme execution

Native RELION Schemer operators, fork, cooperative abort-file polling, explicit resume node, incremental file arrival and reset. Operator-only loops required an explicit abort-file check. No movie processing or scientific dependency invalidation is established.