Platform

Multimodal by design

An open multimodal SciFM unifying text, structures, images and simulation data — ingesting heterogeneous materials inputs and returning actionable predictions.

Platform Internal Testing

Text

Structures

Numerical

Images & fields

Property prediction & UQ

Generative inverse design

Synthesizability & routes

Cross-modal search

Roadmap

SciFM releases

Three open model releases from prototype to final platform deliverable.

V0

M18

50–100M parameters

2–3M records · 150–250k materials

Prototype — shake out modality encoders, fusion, and multi-task heads.

V1

M30

120–200M parameters

6–8M records · 250–350k materials

First stable KG slice — 8–10 prediction tasks, physics-informed generative module.

V2

M42

200–300M parameters

10–15M records · 350–550k materials

Final release — experimental feedback incorporated, open-source platform (D17.3).

Training corpus

Data mix

Multimodal pre-training split across text, graphs, numerics and microstructure images.

40%

Text

30%

Graphs

20%

Numerics

10%

Images

Workflow

Closed loop

From generative proposals through simulation verification to lab validation and knowledge-graph feedback.

Step 1

Propose

SciFM generative engine proposes material candidates from multimodal embeddings.

Step 2

Verify

AI-accelerated surrogates run DFT → phase-field → CFD/FEM at 10³× speed with <5% error.

Step 3

Learn

Validated results flow into the FAIR KG — ontologies, provenance, leakage-safe splits.

Step 4

Synthesize

Top candidates advance to lab validation; experimental data closes the loop.

Architecture

Platform components

Microservices, generative design, trust layers, NLP interface, open releases and multi-scale surrogates.

Unified SciFM hub

Microservices architecture with 15+ REST APIs connecting the KG, SciFM V1/V2, and the simulation pipeline.

Generative inverse design

Physics-informed diffusion/VAE models with multi-objective optimisation — thermodynamic and charge constraints.

Trust & active learning

Deep ensembles, conformal prediction (ECE ≤ 0.05), Data Gate audits, and simulation-gated candidate selection.

NLP interface

Materials-science NLU for querying the KG, launching virtual experiments and steering inverse design — hosted on SimuPort (simuport.com).

Open by design

Apache 2.0 code on GitHub, models on Hugging Face, datasets on Zenodo — model/data cards and 3-year hosting plan.

Multi-scale surrogates

GNN (DFT), PINN (phase-field), ROM (CFD/FEM) surrogates orchestrated in a containerised Kubernetes pipeline.

Infrastructure

Platform scale

Curated data volume, knowledge graph size, model parameters and pre-training compute.

Curated data points

15M+

Knowledge graph entities

10M+

SciFM parameters

200–300M

Pre-training compute

2M GPU-h

Horizon Europe

Project target

Discovery compression, simulation speed-up, generative quality and TRL progression across 48 months.

Discovery

>50%

less time-to-discovery

Simulation

10³×

speed-up at <5% error

Generative quality

>80%

valid, synthesizable proposals

Technology readiness

TRL 1→4

in 48 months

See also the integration diagram on the homepage.

Project updates subscription

Get milestones, publications, events, consortium news and other SimuLingua project information — straight to your inbox.

Project coordination

Get in touch with the SimuLingua project.

FLOWPHYS AS coordinates the Horizon Europe action. Reach out for scientific, consortium or press enquiries across our nine partners.

Project enquiries

HORIZON-RIA · GenAI4EU

1 June 2026May 2030

HORIZON-CL4-INDUSTRY-2025-01-DIGITAL-61

Follow the project on our social networks