Blog

Making AI Execution Predictable: Q&A with Ethan Salehi

Written by Lynx | Oct 6, 2026, 9:05:10 PM

The second installment in our three-part MOSA.ic.AI webinar series with Military Embedded Systems tackled a question that comes up when AI leaves the lab: what happens when a model that performed well in benchmarks has to share a processor with a safety-critical function? Ethan Salehi, Technical Account Manager at Lynx, walked through the engineering decisions that turn an AI workload into something a mission computer can trust, before the audience Q&A dug into specifics on timing, GPU hangs, and what "deterministic AI" really means. We've published the full exchange below.

If you'd like to watch or listen to the full presentation, "Making AI Execution Predictable: Deterministic AI for Safety Critical Systems," register at https://www.lynx.com/mosa.ic.ai-webinar-series.

 

What the Webinar Covered

Salehi opened with a scenario every integrator recognizes: a model performs well in isolation, then moves into a mission computer next to a display and a periodic safety function, and the questions multiply fast. What happens when the AI workload runs late or stops responding? What happens to everything else sharing that processor?

The session walked through the full chain a sensor sample travels: acquisition, time stamping, buffering, preprocessing, inference, post-processing, and a safety gate that decides whether the result is usable. Salehi noted a common mistake: optimizing the neural network does nothing if a frame is already stale by the time it reaches the front of the queue. He covered how ComputeCore and Vulkan SC give an integrator accelerated, analyzable execution paths for CPU and GPU workloads, how TrueCore monitors GPU processing integrity, and how LynxSecure's partition boundaries keep a misbehaving AI workload from stealing time or resources from a higher-criticality function.

Several architecture scenarios were looked at where AI sits relative to safety-critical code, including a lower-criticality AI service feeding a safety gate and a scenario where the AI function itself carries safety responsibility.

Key topics and takeaways included:

  • Bounding the Whole Path, Not Just Inference: Timing budgets have to account for acquisition, buffering, preprocessing, and post-processing, not just the neural network's own execution time.
  • CPU and GPU Isolation Together: Memory protection has to cover both CPU access and device access, including DMA remapping through an IOMMU or SMMU, so one subject can't interfere with another's resources.
  • The Safety Gate Pattern: A dedicated gate checks AI results for freshness, format, and expected range before a mission application takes action, and defines a fallback behavior when a result is rejected or missing.
  • GPU Management and Recovery: TrueCore and the GPU manager detect failures like a hang or lost progress and define how a GPU reset is handled, including which clients are affected.
  • A Workflow From Training to Target: Freezing a deployment baseline, defining a configuration policy with the LynxSecure separation kernel, and testing under realistic concurrent workloads, not just worst-case load.

Read on for the audience questions and Ethan's answers.

 

Audience Q&A

Q: When you say "deterministic AI," what is actually bounded, and how do you establish that bound on hardware?

Ethan: That's a good question. We're aiming to bound the deployed workload's behavior, including the time from acquisition to an acceptable result, its resource demands, and its response to defined failures. That's what deterministic means in this context. It doesn't mean the AI model itself is deterministic. You're looking at this from the safety-critical application's perspective and asking what deterministic behavior that application needs to receive in order to act on it. That's what deterministic execution means here.

 

Q: Can we import our existing ONNX model directly, and what has to be re-verified when we update it?

Ethan: The first step is always to inspect the model against the supported path for the ComputeCore release. Right now the Neural Network Exchange Framework, or NNEF, is being actively worked and supported, and ONNX was supported in an older version of ComputeCore. But those labels alone don't establish compatibility. The number of operators, tensor shapes, precision, custom layers, and conversion options all matter, so we need to identify exactly what has to change in the operation and execution exchange. This is the part I think users and integrators appreciate when we support them through it. We tell them how the model needs to be imported, how it needs to be executed, and what timing, resource, and sharing controls need to be in place for the subject deploying that model.\

 

Q: In AI-aided systems, there's always a question about traceability across lifecycle artifacts. How do you guarantee deterministic traceability across the lifecycle?

Ethan: That's a very good question, and the honest answer is you can't guarantee determinism in the AI model's internal details. The whole point of a neural network is that it trains itself based on the parameters and data it receives. What traceability means here, especially in a mixed-criticality system, is maintaining traceability for the safety-critical components, where data flow and control flow interference actually matter. The integrator has to account for the scenarios where an interference channel could be introduced across those boundaries and define how to mitigate it. That requires the same process you'd follow for standards like CAST or AC21.

 

Q: If the AI-related process doesn't finish in its allocated time, what happens? Does it kill the process, or allocate extra time from future time slots?

Ethan: It depends on your requirements. Giving the AI more time to execute usually doesn't help, because new calculations keep arriving, and those calculations shift and affect every input and output downstream. It really comes down to how your system works and what capabilities you need from the AI. Part of what TrueCore and the GPU manager do is detect when the AI's allotted execution time has passed without producing the expected result, or when the result doesn't match the expected input format for the safety-critical function. That triggers an error that gets fed back to the safety-critical application. That's how we control the deterministic boundary.

 

Q: If a lower-criticality AI job is already running, how can the platform protect a critical display deadline, especially if the GPU hangs?

Ethan: Part of the whole point of TrueCore and the GPU manager is preventing the GPU from entering undefined behavior. Earlier, we talked about GPU reset and where it applies. When that reset happens, we have to account for every client affected by it, including displays. The same failure mode analysis that teams already apply to safety-critical or ARINC-based system assessments needs to apply to AI here too, specifically around what happens when the GPU hangs because the AI is holding onto resources. At some point, the integrator has to control the software so the GPU either gives up the task or resets appropriately, and define what every other application affected by that reset needs to do in response.

 

Q: Boundless architectures aren't new to the industry. Integrated Modular Avionics already defines resource sharing and data flow across an aircraft data network. How does MOSA.ic AI differ from IMA architectures?

Ethan: That's true, but it doesn't fully account for what we're seeing in the technology today. One challenge with older IMA implementations was turning off cores in a multicore environment and falling back to a single core to preserve ARINC 653 time and space partitioning across components. What MOSA.ic AI is trying to do is leverage multicore technology rather than reduce it just to simplify verification against the standard. Specifically with the unikernel approach, we want to preserve time and space partitioning while giving each subject its own resources. That way, when you get to verification and validation, you aren't accounting for every concurrent application running at once by yourself. Part of that burden shifts to the certification of the LynxSecure package and its configuration policy, which can be done on your behalf. Any remaining interference opportunities still need to be mitigated, but those become issues at the application level rather than something you didn't think about until you moved to multicore, which I've seen plenty of in past multicore certification work.

 

Q: Does putting AI in a lower-criticality partition behind a safety gate remove the need to assure the model?

Ethan: No, it doesn't. One of the points we emphasize is that deterministic execution still has to be preserved. The whole point of using ComputeCore alongside TrueCore is to provide an ecosystem where the AI executes in a time-sliced manner within its allocated window. If the AI doesn't behave within that window, the value it provides to your safety application drops and loses credibility fast. We want to get the most out of the AI we can, and that's exactly why deterministic execution matters. The AI we're aiming for with MOSA.ic AI is one that executes with predictable timing, within the space and resources available to its subject or operating system. The feedback that AI receives can then feed back into regenerating or retraining the model through the shift-left containment methods we're building toward, so the model keeps maturing without disrupting the project timeline.

Have questions about integrating AI workloads alongside safety-critical functions? Contact the Lynx team to discuss your system requirements.

 

Continue the MOSA.ic.AI Series

Watch Part 1: Lynx Principal Software Engineer, Ernie Harrison, explores Bridging the AI Deployment Gap with MOSA.ic.AI: Your Trained Model, Deployed for Deterministic Inference in Constrained, Safety-Critical Environments. Watch the recording here.

Join us for Part 3: Watch us explore building End-to-End AI/ML Workflow Blueprints for Safety- and Mission-Critical Systems alongside Ansys, part of Synopsys on Tuesday, October 20 at 2 p.m. ET. Register here.

 

Continue the MOSA.ic.AI Series

Talk with Lynx about bringing trained AI models into embedded systems with deterministic inference and support for safety-critical workloads.

Learn more about MOSA.ic.AI here.