High-throughput data generation
Automated the complete simulation lifecycle, enabling large numbers of physics-based cases to be generated with significantly less manual intervention.
An automated computational workflow for generating high-fidelity engineering datasets, training scientific machine-learning models and validating them against physics-based solvers.
AI for physics has the potential to change how complex engineering and scientific problems are solved. By learning from high-fidelity simulations, scientific machine-learning models can accelerate analysis, support inverse problems, expand design exploration and make computational methods usable at a much larger scale. Over time, this can affect sectors ranging from semiconductors and aerospace to defence, energy, geophysics, healthcare and space exploration.
The central bottleneck is data. In conventional AI, large datasets may already exist or can be collected from digital systems. In scientific machine learning, each data point may represent a full physics-based simulation. Generating one sample can require geometry creation, meshing, solver execution, convergence checks, post-processing and storage. For demanding three-dimensional problems, this makes dataset creation one of the most computationally expensive parts of the entire workflow.
The synthetic-data generation pipeline was developed to address this bottleneck. The objective was to turn simulation data generation from a manually managed sequence of solver runs into a repeatable, scalable and validation-aware computational system.
The workflow begins by defining the parameter space for the engineering problem. This may include geometry, material properties, source conditions, boundary conditions, operating parameters and the physical quantities that need to be predicted. Sampling strategies are then used to generate simulation cases across this design space.
Each case moves through an automated pipeline covering geometry preparation, mesh generation, solver configuration, parallel execution, error handling, convergence monitoring, post-processing and structured data extraction. The system is designed to manage large simulation campaigns while maintaining traceability between input parameters, solver outputs, metadata and validation checks.
The resulting datasets are prepared for scientific machine-learning applications such as surrogate modelling, operator learning, inverse modelling, optimisation and rapid field prediction. Model outputs are then compared with trusted numerical-solver results so that speed improvements do not come at the cost of physical credibility.
Parameter space → automated simulations → structured dataset → scientific ML → solver-based validation
The work resulted in scalable simulation and data-generation workflows developed for international engineering clients, combining high-fidelity solvers, automation, parallel computation and scientific-ML readiness.
Automated the complete simulation lifecycle, enabling large numbers of physics-based cases to be generated with significantly less manual intervention.
Applied advanced algorithms, workflow optimisation and parallel execution strategies to achieve speed improvements of up to 100× in selected data-generation problems.
Produced structured, traceable and validation-aware datasets that could be used for surrogate models, inverse problems, optimisation and other AI-for-physics applications.
I’m always open to conversations about engineering, computation, deeptech products and mentoring people building or learning in these areas.