BITMAIN
Antminer S21+ Hydro Hashboard Repair Service
Antminer S21+ Hydro Hashboard Repair Service
By placing your order you agree to our Terms of Service, Warranty Terms and Refund Policy.
Drop off available in Fort Lauderdale, FL (2141 NE 51st Ct).
Repair guide last verified:
▸How to ship your hardware
Share

An Antminer S21+ Hydro stops the whole machine — all three boards — over a fault that lives on one of them, and the line it prints doing it often names the wrong thing. A healthy S21+ Hydro board reports 95 chips, and any other number is a fault. Both are repaired at component level in the USA, 4-5 business days.
Most Antminer S21+ Hydro hashboard repair jobs that reach our USA-based lab arrive described as a cooling problem, and most of them are not one. The machine stops, the log says it is too hot, and the owner starts on the chiller, the pump and the coolant — while the machine was never hot in the first place. Which of the three boards stopped it is not something the dashboard will tell you, and it is not your job to work out: the miner comes in complete, goes onto our bench loop, and the board is found, repaired at component level and proved under load before the machine goes back.
Technical Diagnostics and Common Issues with Antminer S21+ Hydro
Quick Symptom Checklist
These are lines this model prints on stock firmware, copied as the firmware writes them. Chain numbers, chip counts and temperature figures differ from machine to machine, so match the shape of the line rather than the digits. The machine counts its boards from zero, which makes chain 2 the one you would call the third.
-
ERROR_TEMP_TOO_HIGH: over max tempfollowed bystop_mining: over max temp— and on the line above it, readings of0where the temperatures should be -
got abnormal water temp_max 65409— a number that is not a water temperature at all, printed as if it were one -
Sweep error string = P:1.orJ0:6.orV:1.in the lines around a stop — the letter names the class of the stop, not the fault -
EEPROM error: CRC 1ST REGIONandgot nothing, ending inData load fail for chain 0.— the controller could not read that board's stored identity -
fail to read eeprom by iic, chain: 0, addr: 0andiic write failed— usually arriving in the same minute as the temperature nonsense above -
Chain 0 only find 0 asic, will power off hash board 0withERROR_ASIC_NUM: asic number is not right— a board that came up short of its 95 chips -
ERROR_SOC_INIT: basic init failed!or a boot that ends onsomething error when miner init, restart...and loops there - Not in the log at all: a machine that runs for an hour or two, drops out of the pool, comes back after a restart, and does it again
Model-Specific Patterns We See on Antminer S21+ Hydro
An overheat stop with no heat anywhere in it. The firmware announces ERROR_TEMP_TOO_HIGH: over max temp and shuts the machine down, and the line directly above it reads every temperature on that board as 0 against ceilings of 85 and 95 °C. Nothing was hot. What the machine is saying is that the figures it was given are not figures it can run on, so it stopped rather than guess. It is not a verdict on your cooling, and this is the pattern that costs owners the most time, because the message names the one thing that is not wrong.
A water temperature that cannot exist. got abnormal water temp_max 65409 is the same story wearing a different face — a figure so far outside anything physical that the firmware flags it as abnormal, then stops the machine a moment later anyway. A real coolant problem moves gradually and moves both ends of the loop together. A number like that one came from the board, not from the water.
One board's fault takes all three down. Stock firmware on this model does not run the machine short. A fault confined to a single board does not cost you a third of your rate, it costs you the lot, with two healthy boards sitting idle beside the one that stopped. That is why an S21+ Hydro that has been down for a fortnight is usually a one-board job, and why the arithmetic of repairing it against replacing the machine is not close.
Several different faults print the same way. An i2c line, a chain that finds none of its 95 chips, and a failure to read the board's stored identity often arrive in the same part of the log, and owners read them as one catastrophic event. They are neither one fault nor one cause. There are many ways to end up at those lines, the log separates none of them, and which one it is on your board is bench work — found with the board warm and running under coolant, not read off a screen.
Hardware Notes
| Specification | Details |
|---|---|
| Algorithm and cooling | SHA-256. Liquid-cooled, with no fans anywhere in the machine: each board carries its own cold plate and the three plates share one coolant loop. Bitmain writes the model as S21+ Hyd., and so does the firmware — updated miner type to: Antminer S21+ Hyd.
|
| Miner hashrate | 355 TH/s at roughly 5360 W for the bin the repair documentation describes; the same machine is sold in several bins between 319 and 395 TH/s. One board out is a third of whatever your unit is rated at, and stock firmware does not run the machine on the other two. |
| Hashboards per miner | 3. The log and the interface call them chain 0, chain 1 and chain 2, and the boot line that names all three at once is chain_num 3. |
| Chips per board | 95, so 285 in the machine. The healthy line is chain_asic_num 95, and a good boot prints it three times as Chain[0]: find 95 asic and the same for chains 1 and 2. Any other number on a board is a fault, whether it is 94 or 0. More on what the count means is on our page about ASIC count on the hashboard. |
| How the chips are grouped | 19 groups of 5, which the firmware states as chain_domain_num 19 and domain_asic_num 5. Domain is the firmware's own word for a group, and it is the word to look for in your log. |
| Chip family |
BM1370, the same family used across the 21 series in several letter-suffixed versions. The suffix is not documented by anyone, and the working rule is to match the marking already on the board rather than to reason from the model name. |
| Hashboard part number | The interface header shows it as Antminer H6HB70701, and the boot log prints the configuration it loaded for it — load machine H6HB70704 conf. Those two strings differ on a perfectly normal machine and both belong to the same board; a build stamped 70701, 70702 or 70704 is the same hashboard. H6 at the front is Bitmain's mark for a liquid-cooled 21-series board, and it is what separates this board from the air-cooled S21+, whose number ends in the same digits behind a different prefix. |
| Firmware stop limits | The log prints its own ceilings beside every reading: (max 85) for the board and (max 95) for the chips. It is worth reading those two figures next to the numbers in front of them, because that is where a stop with zeros in it gives itself away. |
| Power supply | APW111721 series. On a healthy boot the log identifies it and calibrates it; when the controller cannot get an answer it prints get power version failed and retries on an older protocol, and on its own that line is not a verdict on anything. |
What the Log Is Telling You
None of this asks anything of you. It is here so you know what state your machine is in before you decide to ship it.
| What you see | What it usually means | What we do with it |
|---|---|---|
ERROR_TEMP_TOO_HIGH: over max temp with zeros in the readings on the line above |
Not a heat problem. The machine was given figures it could not run on and stopped rather than guess. | We read the numbers against the ceilings printed beside them, then work the board warm and under coolant. A stop shaped like this points at the board, and we tell you so before you spend anything on the loop. |
got abnormal water temp_max 65409 |
A figure that no coolant loop can produce. Real cooling faults move gradually and move both ends of the loop together. | We treat it as coming from the board rather than from the water, and prove that on the bench under coolant instead of replacing cooling hardware to find out. |
EEPROM error: CRC 1ST REGION, got nothing, Data load fail for chain 0.
|
The controller could not read that board's stored identity, so it has nothing to configure the chain from. Several unrelated faults end here. | We establish which of them it is on your board, repair it, and confirm the board is read correctly on a cold start and again warm. |
Chain 0 only find 0 asic, will power off hash board 0 and ERROR_ASIC_NUM: asic number is not right
|
That board reported fewer than 95 chips, so the firmware shut it down. Zero and a count a few short are the same class of fault, not different severities. | We bring the chain up on the bench under coolant and find the point where the count breaks, then repair at component level and read the full 95 back. |
Sweep error string = P:1., J0:6. or V:1.
|
An index, not a diagnosis. The letter names a class of stop, and the lines around it say what actually happened. | We use the letter to find the right part of the log and then read the log, because on this model the same letter arrives for faults with quite different cures. |
Diagnostics Focus
Believe the count, question the temperature. On this board the chip count is a reliable witness and a temperature reading is not: 95 or a fault, with nothing in between to interpret. A temperature, by contrast, can be missing, frozen or impossible while the machine is stone cold. So the order of work is the opposite of what the log suggests — the readings are established as real or not real first, and only a board producing honest readings is judged on what those readings say.
Warm, wet and for long enough. A board that passes cold proves very little here. The faults that bring these machines in are the intermittent kind: a chain that holds 95 chips at boot and loses them once the machine is hot, a stop that only arrives after twenty minutes under load. Every measurement is taken with coolant flowing and the board at working temperature, and a repaired board goes back only after it has held its count and its readings under real load.
Our Professional Repair Process
Gotchas
The coolant comes out before it ships, and the ports go back on. A machine that travels with liquid still in it usually arrives with some of that liquid somewhere it should not be, and a board that has been wet and then powered is scrap rather than a repair. Drain it, cap the inlet and the outlet, and it travels without incident. This is the one thing we ask of you before the machine leaves.
The cold plate is bolted down, not soldered — which is the whole reason this board is repairable. Owners are told a hydro board is a sealed unit to be replaced rather than fixed, and it is not: the plate comes off, the chips underneath are reachable, and the board goes back together. What it does mean is that the compound under that plate is renewed every single time it is opened. A plate put back onto its old compound runs hotter afterwards than it did before, which is how a board can reach us in worse condition than it was in when somebody else opened it.
Typical Service Scenario
Months spent on a cooling loop that was never at fault. The machine stops on temperature, so the owner works through the obvious list — flushes the loop, changes the coolant, checks the pump, eventually replaces cooling hardware — and the stops keep coming back at the same interval. The fault was never in the water, and everything that was done to the loop was done to a healthy loop.
A hosted machine handed back. These units run in hosted halls on three-phase power, and when one drops off the floor the site pulls it and returns it to its owner, who now has a 355 TH/s machine that will not start and no local service option. The factory route for a liquid-cooled Antminer is measured in weeks or months, in both directions.
An opened machine that got worse. Somebody lifted a cold plate to look, put it back on what was already there, and the machine came back with temperatures higher than before and a new stop pattern. That is a repairable board that has been made harder to repair, not a ruined one.
What Happens After Intake
The machine comes onto our bench loop whole: intake, incoming inspection of the coolant path and of all three boards, then diagnostics with the boards running under coolant rather than dry. Component-level repair follows, then cleaning and a freshly printed thermal interface under every plate that was lifted. Board-level validation runs on our STASIC and ASIC REPAIR fixtures, and the machine then runs under load for a minimum of one hour before it goes back. Extended burn-in beyond that hour is a separate add-on.
Diagnostics and Validation Equipment
Our ASIC REPAIR and STASIC fixtures read the board's chip inventory back the way the miner does, so a pass on this model means the fixture found all 95 chips and the board came through its checks clean — not merely that nothing obvious was wrong. The verdict that counts is still the machine's own: the repaired board back in its slot, in the loop, holding its count and its share of the rate under load for at least an hour.

All repairs are performed personally by Alex, who has been servicing ASIC miners since 2019. See credentials →
Contact our repair team today and get your miner back to full power.