35B Local LLM Workload Triggers DDR4 ECC Errors in Aging Workstation
A local AI user running a 35B language model on an older workstation has reported repeated DDR4 ECC memory errors during sustained inference, with problems appearing on 2 different modules at separate points in the testing period. The heavier workload appears to have exposed a stability issue that had not appeared during lighter use.
The case originated from a post in the LocalLLM community on Reddit, where the user described running a 2019 Lenovo ThinkStation P920 with dual Xeon processors, 128 GB of ECC DDR4 memory using 8 GB modules, and an NVIDIA GeForce RTX 5070 Ti with 16 GB of VRAM. The system had reportedly handled models around 20B parameters and smaller for approximately 6 months without memory related problems.
The situation changed when the user moved to Ornith 1.5 35B for a database workload that smaller models were struggling to complete. Because the larger model exceeded the available GPU memory, part of the workload spilled into system RAM.
The first episode occurred during the initial extended run. After approximately 90 minutes, the workstation crashed and the user found memory errors in the system logs. The model was restarted and the system crashed again within approximately 20 minutes. On a third attempt, the user monitored the logs in real time and reported ECC errors accumulating on the module installed in slot 8 after around 10 minutes. That module was then removed, after which the same workload reportedly ran without further problems at that stage.
The user initially treated the incident as an isolated problem with that module and continued using Ornith 1.5 35B because its database performance was considerably better than the smaller models previously tested.
The second episode did not happen immediately afterward. The workstation continued running the 35B model heavily on a daily basis for another approximately 8 to 9 days. Only after that period did the user check the logs again and report errors associated with a second memory module.
This creates 2 distinct points in the timeline: one module began reporting errors during the first prolonged 35B session and was removed, while another module reportedly began showing errors more than a week later after continued heavy daily use. ECC reporting confirms that corrupted memory data was detected, but it does not establish permanent physical damage to either module. The original post also does not specify whether the logged events were corrected or uncorrectable errors.
The timing remains notable because the system had reportedly spent approximately 6 months running smaller models without similar issues. The larger 35B workload depended more heavily on system memory, keeping more of the installed RAM active for extended periods and potentially exposing instability that lighter workloads had not revealed.
Temperature is another unresolved variable. The user reported CPUs staying below approximately 80°C and the RTX 5070 Ti generally below 65°C, with additional case airflow installed. No direct DIMM temperature measurements were provided.
All installed memory was also described as matched and originating from the same batch. That becomes more relevant once a second module begins producing errors days later, although the available information is still insufficient to identify a common underlying cause.
This is particularly relevant as older workstation platforms become increasingly attractive for local AI. Large DDR4 ECC configurations can provide substantial memory capacity at comparatively low cost, but sustained inference workloads can exercise those systems differently from conventional desktop applications.
The GMKtec EVO X5 Pro with 192 GB of unified memory for models exceeding 300B parameters, reflecting how system memory capacity and sustained memory access are becoming increasingly important as users run larger models locally.
The timeline makes this case more interesting than a single crash, but it also makes the "LLM destroyed RAM in 90 minutes" interpretation difficult to support. The first module began producing identifiable ECC errors during the initial extended workload and was removed. The system then continued operating under heavy 35B inference for another 8 to 9 days before a second module began reporting errors.
What the report shows most clearly is that sustained local AI use coincided with memory instability appearing across 2 different modules over time. It does not establish that either DIMM was physically damaged by the model.
For anyone repurposing aging workstation hardware for local AI, the case is still worth paying attention to. Workloads that continuously involve large amounts of system RAM can expose problems that lighter everyday use may never reveal.
Have you seen memory errors appear only after several days of sustained local AI workloads on older DDR4 hardware?
