V0
M18
50–100M parameters
2–3M records · 150–250k materials
Prototype — shake out modality encoders, fusion, and multi-task heads.
Platform
An open multimodal SciFM unifying text, structures, images and simulation data — ingesting heterogeneous materials inputs and returning actionable predictions.
Text
Structures
Numerical
Images & fields
Property prediction & UQ
Generative inverse design
Synthesizability & routes
Cross-modal search
Roadmap
Three open model releases from prototype to final platform deliverable.
V0
M18
50–100M parameters
2–3M records · 150–250k materials
Prototype — shake out modality encoders, fusion, and multi-task heads.
V1
M30
120–200M parameters
6–8M records · 250–350k materials
First stable KG slice — 8–10 prediction tasks, physics-informed generative module.
V2
M42
200–300M parameters
10–15M records · 350–550k materials
Final release — experimental feedback incorporated, open-source platform (D17.3).
Training corpus
Multimodal pre-training split across text, graphs, numerics and microstructure images.
40%
Text
30%
Graphs
20%
Numerics
10%
Images
Workflow
From generative proposals through simulation verification to lab validation and knowledge-graph feedback.
Step 1
SciFM generative engine proposes material candidates from multimodal embeddings.
Step 2
AI-accelerated surrogates run DFT → phase-field → CFD/FEM at 10³× speed with <5% error.
Step 3
Validated results flow into the FAIR KG — ontologies, provenance, leakage-safe splits.
Step 4
Top candidates advance to lab validation; experimental data closes the loop.
Architecture
Microservices, generative design, trust layers, NLP interface, open releases and multi-scale surrogates.
Microservices architecture with 15+ REST APIs connecting the KG, SciFM V1/V2, and the simulation pipeline.
Physics-informed diffusion/VAE models with multi-objective optimisation — thermodynamic and charge constraints.
Deep ensembles, conformal prediction (ECE ≤ 0.05), Data Gate audits, and simulation-gated candidate selection.
Materials-science NLU for querying the KG, launching virtual experiments and steering inverse design — hosted on SimuPort (simuport.com).
Apache 2.0 code on GitHub, models on Hugging Face, datasets on Zenodo — model/data cards and 3-year hosting plan.
GNN (DFT), PINN (phase-field), ROM (CFD/FEM) surrogates orchestrated in a containerised Kubernetes pipeline.
Infrastructure
Curated data volume, knowledge graph size, model parameters and pre-training compute.
Curated data points
15M+
Knowledge graph entities
10M+
SciFM parameters
200–300M
Pre-training compute
2M GPU-h
Horizon Europe
Discovery compression, simulation speed-up, generative quality and TRL progression across 48 months.
Discovery
>50%
less time-to-discovery
Simulation
10³×
speed-up at <5% error
Generative quality
>80%
valid, synthesizable proposals
Technology readiness
TRL 1→4
in 48 months
See also the integration diagram on the homepage.
Project updates subscription
Get milestones, publications, events, consortium news and other SimuLingua project information — straight to your inbox.
Project coordination
FLOWPHYS AS coordinates the Horizon Europe action. Reach out for scientific, consortium or press enquiries across our nine partners.
FLOWPHYS AS
Per Kjellgren · Project Coordinator
Oslo, Norway
View consortium partnersper.kjellgren@flowphys.com
General enquiries
contact@simulingua.eu
+47 40621185