Rule engine
Warn and alarm thresholds, slope checks and simple AND/OR logic per failure. Transparent, and enough to run the whole loop.
Example: vibration above 2.5× baseline, still rising, with motor temperature climbing, means bearing wear.
NASA HUNCH · Kennedy Space Center · LLASO Project 4
Robots on the Moon work for months with no technician nearby. Luminary reads their telemetry, finds the part that is wearing out, estimates how long it has left, and sends the robot for repair while a spare takes over its job.
Crewed lunar missions last 30 to 60 days. After the crew leaves, the robots keep working on their own for months. A drive motor that quits in the wrong place strands a rover that a mission was counting on.
Asking Earth what to do does not work. A signal takes about 2.5 seconds each way, practical delays run to minutes, and a robot behind a crater rim has no link at all. So the decision has to be made on the Moon, before the failure, by software.
Every robot streams motor, wheel, battery, thermal and context data. Luminary runs the same seven steps on all of it, continuously.
Clean the six telemetry streams and compare each signal against the robot's own normal for the same load and operating mode.
Score how far current behavior is from normal, from 0 to 1.
Name the likely failure mode and the part responsible, with a confidence, and explain it in plain language.
Estimate time to failure for the degrading part, with a confidence interval.
Combine severity, the robot's role, urgency and confidence into one criticality score.
Ask the robot to run its own self-check. It can confirm the flag or reject it.
Write a work order, queue it for the repair garage, and launch a standby robot into the role.
An illustration of one failing front-left drive motor with a bearing wearing out. Vibration climbs as failure approaches, and Luminary raises flags at the thresholds from the project spec. The curve is illustrative, not recorded data.
All four drive motors are within their normal range. Nothing to do.
Each row is one simulated robot whose motor bearing wore out. The chart shows how many minutes before failure each method raised a confirmed flag. Further right is earlier. Both use the settings the Luminary app runs with.
The rules are tuned to stay quiet: a motor must run at 1.8 times its normal vibration, rising, for 5 seconds in a row before they fire. That kept them to zero false alarms, but they fire late, and every robot got less than the 15-minute warning we target. Machine learning warned earlier on all 15 robots, with at least 21 minutes on every one.
On top of that it estimates time to failure, within about 8 minutes on average, and names the likely cause. The hard-limit rules still run beside it as a safety net.
Everything in the comparison above comes from simulated robots, scored the same way for both methods.
A telemetry generator produced 300 robot runs of about an hour each, one reading per second, using a fixed random seed so the run can be repeated. 117 of the runs had a motor bearing slowly wearing out until it failed. The generator also adds sensor spikes and about 1% missing readings, the way real sensors misbehave.
The runs were split by robot into training, validation and test groups, so no robot the models were scored on was ever used to train them. The test group has 45 robots: 15 with the failing bearing and 30 healthy.
The data is cut into short sliding windows, one every 10 seconds, and summarized into features such as averages, trends and spread. A random forest learns to spot bearing wear, a second one estimates time to failure, and an Isolation Forest looks for anything unusual.
The machine learning flags a window when it is at least 80% sure. The hard-limit rules flag a motor when its vibration is above 1.8 times normal and rising, and only after that holds for 5 readings in a row. These are the same settings the Luminary app uses.
For each failing robot, warning time is how many minutes of life it had left when the method first raised a flag. For each healthy robot, we counted whether the method flagged it at any point in its hour.
Luminary recognizes these failure modes and the telemetry pattern that gives each one away.
| Failure | Where | What the sensors show | Warning time |
|---|---|---|---|
| Motor bearing wear | Motor | Vibration and ultrasonic noise rise, temperature creeps up | 15 min to hours |
| Winding or rotor fault | Motor | Current sidebands, higher current, higher temperature | Minutes to tens of minutes |
| Drive-wheel motor stall | Wheel | Acceleration spike, current spike then drop, wheel imbalance | Seconds to minutes |
| Gear backlash or joint wear | Joint | Position error and backlash grow, joint current rises | Hours to days |
| Encoder or sensor fault | Joint / IMU | More dropouts, gyro or heading bias, sudden residual jump | Immediate |
| Battery capacity fade | Battery | State of health falls, internal resistance rises | Many cycles |
| Weak or failing cell | Battery | Cell-voltage spread widens, local overheating | Hours to cycles |
| Thermal-control fault | Thermal | Bay temperature drifts, heater duty rises | Minutes to hours |
| Dust or seal degradation | Mechanical | Dust-ingress proxy rises, bearing signature speeds up | Long horizon |
The same inputs and outputs work at every tier, so a simple version always runs and the smarter versions layer on top.
Warn and alarm thresholds, slope checks and simple AND/OR logic per failure. Transparent, and enough to run the whole loop.
Example: vibration above 2.5× baseline, still rising, with motor temperature climbing, means bearing wear.
Isolation Forest for anomaly scoring, Random Forest or gradient boosting for failure mode and time to failure, trained on engineered features from 30 to 60 second windows.
Small LSTM autoencoders and temporal CNNs that learn normal operation and flag what does not fit. Kept tiny enough to run on the robot.
The hard-limit rules, such as over-temperature, always run alongside the ML models. If a model misses something dangerous, the rules still raise it.
When a robot is confirmed to be failing, Luminary sends a healthy standby into its role before the failing robot leaves.
The repair garage has limited bays. Luminary ranks every work order by a criticality score and fills the bays from the top.
Severity counts for 35%, how important the robot's role is right now for 25%, how soon it will fail for 30%, and the AI's confidence for 10%. The weights are configurable.
Critical jobs never wait behind lower ones, and long-waiting jobs gain priority so none are starved. Robots waiting for a bay are slowed down or parked safely to stretch their remaining life.
If the robot's own self-check rejects a critical flag, the work order stays and is marked for review. Safety wins over the override.
The acceptance tests the system is measured against.
A live dashboard that runs the whole fleet simulation and shows what the AI sees.
A 3D model of the selected robot with a status light on each wheel motor and a health percentage for every location.
Every robot at a glance, with live data streams per location.
Signal history and trends for any part, using Mann-Kendall, Sen's slope, CUSUM and z-score detection.
Replacement scheduling based on mean time between failures.
A queue of flags with a locked-alert pop-up for critical ones.
Ask which robots are at risk this week and get an answer written from the risk records.