Summary: NASA and IBM released a multimodal model of the lunar surface. Weights, data and adaptation code are public; pretraining code is missing. The cited “up to 23 percent” is not a universal accuracy score, but a summary of task-specific tests.
Two million image tiles at multiple resolutions
The model was trained primarily on 17 years of Lunar Reconnaissance Orbiter data. NASA lists roughly two million image tiles: more than one million camera images at one-metre resolution and nearly 964,000 multispectral images at 100 metres per pixel. Data from GRAIL, Lunar Prospector and Japan's SELENE mission were also included.
Its core uses TerraMind, an Earth-observation model from IBM and the European Space Agency. It brings measurements from different instruments, viewing angles and resolutions into a shared representation. Researchers can then adapt it to specific tasks with comparatively small labelled datasets.
NASA prioritises three applications: detecting small craters, outlining young volcanic structures and estimating the stability of potential ice deposits near the lunar poles. The latter can guide selection of scientifically and technically interesting sites. A model prediction is not direct evidence of water ice, however, and cannot replace measurements on the ground.
What the benchmarks actually show
The broad vendor claim of “up to 23 percent better” hides substantial differences. IBM reports a 22 percent reduction in error for ice prospectivity compared with a specialised SwinV2 model. For crater detection, the model matched the baseline on one-metre imagery; at 100 metres per pixel, it performed nearly 19 percent better with half the training data. Its lead in segmenting irregular mare patches was three percent.
The organisations involved produced these results; independent reproduction is still pending. The percentages also use different tasks and reference metrics, so they cannot be compared directly.
Open, but not fully reproducible
Model weights, configurations, benchmark datasets, and code for fine-tuning and inference are available on Hugging Face and GitHub; the repository uses Apache 2.0. Its README also states explicitly that pretraining code is not included. External teams can inspect and adapt the model, but cannot recreate its complete training process from the repository alone.
Pandorex assessment: The practical value lies less in a single headline score than in combining highly varied lunar measurements. Independent replication, clearly documented compute requirements and tests beyond the three published tasks are now essential for robust scientific use.
