Luminary

NASA HUNCH · Kennedy Space Center · LLASO Project 4

Luminary sees a lunar robot fail before it does.

Robots on the Moon work for months with no technician nearby. Luminary reads their telemetry, finds the part that is wearing out, estimates how long it has left, and sends the robot for repair while a spare takes over its job.

Loading 3D model…
Drag to rotate · scroll to zoom
Median warning before motor failure
15 min or more
Mission gaps in a full simulated cycle
0
Earth in the loop for urgent faults
None
Robot types covered
4

Nobody is there to fix it

Crewed lunar missions last 30 to 60 days. After the crew leaves, the robots keep working on their own for months. A drive motor that quits in the wrong place strands a rover that a mission was counting on.

Asking Earth what to do does not work. A signal takes about 2.5 seconds each way, practical delays run to minutes, and a robot behind a crater rim has no link at all. So the decision has to be made on the Moon, before the failure, by software.

One loop, from sensor reading to repair order

Every robot streams motor, wheel, battery, thermal and context data. Luminary runs the same seven steps on all of it, continuously.

  1. Ingest

    Clean the six telemetry streams and compare each signal against the robot's own normal for the same load and operating mode.

  2. Detect

    Score how far current behavior is from normal, from 0 to 1.

  3. Diagnose

    Name the likely failure mode and the part responsible, with a confidence, and explain it in plain language.

  4. Predict

    Estimate time to failure for the degrading part, with a confidence interval.

  5. Prioritize

    Combine severity, the robot's role, urgency and confidence into one criticality score.

  6. Confirm

    Ask the robot to run its own self-check. It can confirm the flag or reject it.

  7. Dispatch

    Write a work order, queue it for the repair garage, and launch a standby robot into the role.

Drag time forward and watch the warning arrive

An illustration of one failing front-left drive motor with a bearing wearing out. Vibration climbs as failure approaches, and Luminary raises flags at the thresholds from the project spec. The curve is illustrative, not recorded data.

Normal
Vibration vs. baseline
1.0×
Drive-motor health index
100
Time left
120 min
Likely cause
None found

All four drive motors are within their normal range. Nothing to do.

A hard limit that stays quiet warns too late

Each row is one simulated robot whose motor bearing wore out. The chart shows how many minutes before failure each method raised a confirmed flag. Further right is earlier. Both use the settings the Luminary app runs with.

Synthetic data, 15 held-out faulty robots and 30 healthy ones, from the project's own models. Early result, not a guarantee.
Average warning, machine learning
47 min
Average warning, hard-limit rules
7 min
Healthy robots falsely flagged in an hour, machine learning
4 of 30
Healthy robots falsely flagged in an hour, rules
0 of 30

The rules are tuned to stay quiet: a motor must run at 1.8 times its normal vibration, rising, for 5 seconds in a row before they fire. That kept them to zero false alarms, but they fire late, and every robot got less than the 15-minute warning we target. Machine learning warned earlier on all 15 robots, with at least 21 minutes on every one.

How it ignores false alarms

  • Sensor dropouts are skipped, and a single-tick spike is ignored because a flag has to hold for 5 seconds.
  • Each reading is compared with that robot's own recent baseline, not a fixed number.
  • The model has to be at least 80% sure before it raises a flag.

On top of that it estimates time to failure, within about 8 minutes on average, and names the likely cause. The hard-limit rules still run beside it as a safety net.

How we tested it

Everything in the comparison above comes from simulated robots, scored the same way for both methods.

  1. Simulate the robots

    A telemetry generator produced 300 robot runs of about an hour each, one reading per second, using a fixed random seed so the run can be repeated. 117 of the runs had a motor bearing slowly wearing out until it failed. The generator also adds sensor spikes and about 1% missing readings, the way real sensors misbehave.

  2. Keep test robots separate

    The runs were split by robot into training, validation and test groups, so no robot the models were scored on was ever used to train them. The test group has 45 robots: 15 with the failing bearing and 30 healthy.

  3. Train the models

    The data is cut into short sliding windows, one every 10 seconds, and summarized into features such as averages, trends and spread. A random forest learns to spot bearing wear, a second one estimates time to failure, and an Isolation Forest looks for anything unusual.

  4. Run both methods as the app runs them

    The machine learning flags a window when it is at least 80% sure. The hard-limit rules flag a motor when its vibration is above 1.8 times normal and rising, and only after that holds for 5 readings in a row. These are the same settings the Luminary app uses.

  5. Measure warning time and false alarms

    For each failing robot, warning time is how many minutes of life it had left when the method first raised a flag. For each healthy robot, we counted whether the method flagged it at any point in its hour.

Nine ways a lunar robot can break, and what each one looks like

Luminary recognizes these failure modes and the telemetry pattern that gives each one away.

FailureWhereWhat the sensors showWarning time
Motor bearing wearMotorVibration and ultrasonic noise rise, temperature creeps up15 min to hours
Winding or rotor faultMotorCurrent sidebands, higher current, higher temperatureMinutes to tens of minutes
Drive-wheel motor stallWheelAcceleration spike, current spike then drop, wheel imbalanceSeconds to minutes
Gear backlash or joint wearJointPosition error and backlash grow, joint current risesHours to days
Encoder or sensor faultJoint / IMUMore dropouts, gyro or heading bias, sudden residual jumpImmediate
Battery capacity fadeBatteryState of health falls, internal resistance risesMany cycles
Weak or failing cellBatteryCell-voltage spread widens, local overheatingHours to cycles
Thermal-control faultThermalBay temperature drifts, heater duty risesMinutes to hours
Dust or seal degradationMechanicalDust-ingress proxy rises, bearing signature speeds upLong horizon

Rules set the floor. Machine learning adds earliness.

The same inputs and outputs work at every tier, so a simple version always runs and the smarter versions layer on top.

Rule engine

Warn and alarm thresholds, slope checks and simple AND/OR logic per failure. Transparent, and enough to run the whole loop.

Example: vibration above 2.5× baseline, still rising, with motor temperature climbing, means bearing wear.

Classical ML

Isolation Forest for anomaly scoring, Random Forest or gradient boosting for failure mode and time to failure, trained on engineered features from 30 to 60 second windows.

Deep learning

Small LSTM autoencoders and temporal CNNs that learn normal operation and flag what does not fit. Kept tiny enough to run on the robot.

The hard-limit rules, such as over-temperature, always run alongside the ML models. If a model misses something dangerous, the rules still raise it.

A repair never leaves a job uncovered

When a robot is confirmed to be failing, Luminary sends a healthy standby into its role before the failing robot leaves.

  1. ActiveWorking its role
  2. Dispatch pendingFlag confirmed, repair bay planned
  3. Handing offStandby is on its way and takes over
  4. In repairAt the Repair Garage
  5. Standby or activePassed its self-test, back in the pool

Who gets a repair bay first

The repair garage has limited bays. Luminary ranks every work order by a criticality score and fills the bays from the top.

Severity counts for 35%, how important the robot's role is right now for 25%, how soon it will fail for 30%, and the AI's confidence for 10%. The weights are configurable.

When several robots fail at once

Critical jobs never wait behind lower ones, and long-waiting jobs gain priority so none are starved. Robots waiting for a bay are slowed down or parked safely to stretch their remaining life.

If the robot's own self-check rejects a critical flag, the work order stays and is marked for review. Safety wins over the override.

How we know it works

The acceptance tests the system is measured against.

The Luminary app

A live dashboard that runs the whole fleet simulation and shows what the AI sees.

Digital twin

A 3D model of the selected robot with a status light on each wheel motor and a health percentage for every location.

Fleet

Every robot at a glance, with live data streams per location.

Sensor detail

Signal history and trends for any part, using Mann-Kendall, Sen's slope, CUSUM and z-score detection.

Maintenance

Replacement scheduling based on mean time between failures.

Alerts

A queue of flags with a locked-alert pop-up for critical ones.

AI analyst

Ask which robots are at risk this week and get an answer written from the risk records.