Feasibility of running the ET-SoC-1 without its heatsink
No: a lab card cannot run without its heatsink at 600 MHz, the clock three of the four cards run at, nor at 100 MHz, the lowest clock its firmware can set: a slower clock removes little of the heat, which is mostly leakage and power no minion clock drives. Taping a thermocouple to the heatsink base, heatsink on, is feasible, cheap and safe with the card off; pointing a sensor at the bare chip is not. The proposal was to measure the chip's temperature with a sensor of our own, either taped to the heatsink or pointed straight at the chip with the heatsink off. This page works out whether a card can run bare in the lab, what a sensor would see if it did, and which ways of adding a physical temperature are safe. Section 5, added at the owner's request, works out how low the clock can go, what a slow card draws, whether it holds bare, and what a thermal camera would see of where the computation runs.
The proposal, read two ways
Taped to the heatsink: feasible, cheap, safe with the card off.
Soldered into a groove in the lid: only with the lab's consent, at moderate risk.
Either logs a physical temperature independent of the host and its whole-degree sensor mean; the lid groove is where Intel measures case temperature. Each measures one point, not a map (options 2 and 3 in section 6).
Not as a running setup, at any clock the firmware can set.
The card cannot hold a temperature without its heatsink (section 2). In the model, bare aifoundry2 passes 90 °C after a cold power-on at idle in still air (card 1 in ), and with a fan in model (section 3). A single frame shows a lid at nearly one temperature (section 4). At 100 MHz, with a fan, a short look with lock-in is possible in some power-ons, behind an automatic power cut and after the measurement plan (section 5). A shire-scale view of a running chip needs a delidded die on a card set aside for it.
The answers
- Can a card run without its heatsink?
- No, not at 600 MHz, the clock three of the four lab cards run at. With its heatsink a card sheds heat at . Bare, the package can lose heat only through its lid into the air and through its balls into the board: an estimated 4–12 °C/W in still air and 2–7 °C/W with a fan model, a range that holds the vendors' tables for AMD's lidded 40–47.5 mm packages and Microchip's 35 mm MPF500T with no heatsink (5.4–7.9 °C/W in still air, 2.7–5.8 at 1–2.5 m/s) [1, 2]. The card's idle power (mostly the chip's) also rises as it warms, by 0.50–1.02 W for each degree at 70–80 °C on those three cards fitted, so every degree of heating brings more heat. A temperature is stable only while the thermal resistance times that slope stays below 1, which on these cards needs °C/W or less model. Bare, the die keeps heating: From a cold power-on, the model puts bare aifoundry2 in still air past 90 °C after at idle. The host needs an assumed 30–90 s to boot, by which that die is at ; from then a random-data matmul leaves before the die reaches 90 °C model, boot assumed. Nothing has been seen to stop it: at 600 MHz no firmware limit has been seen to act (card 1's build has a 75 W alarm, read from source and untested: section 2.2), short of shutting the whole host down nothing switches the card off, and Esperanto publishes no junction limit or thermal trip [3, 4, 5]. In published bare runs, a processor that throttled itself kept running, and two burned out within a second, one with no protection and one whose board-level protection reacted too slowly [6]; Intel requires a tripped Atom 330's supply to be turned off within 500 ms [7]. The one marginal case is aifoundry1 card 0 idling at 300 MHz, and any kernel lifts it to 600 MHz; on 25 September its release 1.4.1 firmware did step it back to 300 MHz after it read 115–117 °C, by a path not established [R3]. Sections 2–3, 7.1
- Would a much slower clock, 100 or even 10 MHz, let it run bare?
- No, not as a running setup. The firmware can set 100 MHz, six times slower; nothing lower without a new boot loader or a debugger from source. At 100 MHz with the voltages left as they are, aifoundry2's idle at a 60 °C die falls by only from (the measurement plan's model, which allows other forms of the leakage law, gives on aifoundry3), since most of it is leakage and power no minion clock drives; at the firmware's lowest voltages, with the NoC slowed too, it is , and 10 MHz would take only more model. To settle below 85 °C there, a bare card needs a thermal resistance of or less. The bare package has 4–12 °C/W in still air, where no draw settles, and 2–7 °C/W with a fan, where of draws settle on aifoundry2 and aifoundry3, a share that mostly reflects that assumed range model. A 3–6 m/s blower assumed would hold an idle temperature in of draws, but the measurement plan's model brings only of cold power-ons through its gate, which asks for 80% model, untested. A slow clock buys little time: the card idles at 600 MHz while the host boots and the low point is set, so a bare aifoundry2 at 100 MHz then has before the plan's 75 °C stop model, boot assumed. Section 5
- What would a sensor pointed at the chip see?
- In a single frame, one temperature. With the heatsink off, the lid, and the die beneath it, spread heat sideways over mm model, about the width of the die, while the shires sit 3.7 mm apart: a 1 W hot shire would lift the lid by only about °C over a broad patch model. Switching a pattern of work against its complement and averaging the images in step with it (lock-in) does bring patterns out of a taped lid: at 100 MHz and the lowest voltages, a single shire and the coarse patterns within about 10 s and a checkerboard in about model. That fits the minute or so a fan leaves in the better power-ons, not still air (section 5.4). The lid is plated metal, which an infrared camera reads low and mostly as reflections of the room unless it is taped: a study of processor thermography put a metal package's emissivity at about 0.01 and covered the surface with masking tape of emissivity 0.92 [8]. The lab's planned camera sees 8–14 µm, where sapphire, fused silica and Czochralski silicon windows do not transmit well and fluorinated coolants likely absorb; a view of the die in that band would need a germanium, zinc-selenide, zinc-sulphide, chalcogenide-glass or diamond window over a coolant that transmits there. A thin film of mineral oil is the candidate, untested, and it has one absorption band at 13.9 µm, inside that range inference. Section 4
- What should we do instead?
- Keep the heatsink on. For a physical temperature that does not depend on the host, tape a fine thermocouple to the heatsink base or fins (half an hour on site with the card off, $30–100 estimate), never between the lid and the heatsink: AMD warns that one there may add stress and leaves the paste thicker or uneven [1]. For the case temperature itself, and only with the lab's consent and at moderate risk, use Intel's method: a 36-gauge type T thermocouple soldered into a groove cut in the lid's centre [9]; calibrated, it reads the lid to about °C model. Put emissivity tape wherever the camera looks. The natural moment is the on-site check of aifoundry1 card 0's cooling that the lab-problems page already requests (SH1). For where the heat goes, use the thermal-camera plan's view of the back of the board with the on-die voltage map. To see where the work runs, the chip's own 34 sensors are the safe camera: with the heatsink on they see the same contrast between neighbouring shires as with it off model. On the stock firmware the host gets one usable sensor, the I/O shire's, which with lock-in and a slow dither of the chip's power sees a block of shires beside it; a map of every shire needs a rebuilt, signed boot loader, which the cards have never run model, inference (section 5.5). A shire-scale thermal image needs a delidded card that may be lost, an infrared-window cooler and a mid-wave camera, as in the published setups: weeks and thousands of dollars estimate. Section 6
The owner's request (29 September 2026, 17:30 PDT), verbatim
“There was a proposal to install a temperature sensor by taping a heatsink and having a temperature sensor directly pointed at the chip. Can you study the feasibility of this? Can you actually run one of these cards without the heatsink in the lab? Make this a report. Cross-link it to the other reports. I think I had a spaceship report about doing the thermal camera. Make a report called something like ‘Feasibility of Esperanto Without Heatsink’ and cross-link it appropriately.”
Terms and labels used on this page
θ (theta): thermal resistance, the temperature rise in °C per watt of heat. θJA runs from the junction (the transistors) to the ambient air, θJC from the junction to the top of the lid (the case), θJB from the junction into the board. The lid is the metal plate glued over the die; the paste between lid and heatsink is TIM2. Loop gain: θ times the rise of power per degree, dP/dT; below 1 a temperature can settle, at 1 or above it cannot. The Foster chain is the six-stage thermal model fitted to aifoundry2 with its heatsink. The decay length is the distance over which a hot spot's excess temperature falls by a factor of e (2.7) sideways. LWIR and MWIR: long-wave (8–14 µm) and mid-wave (3–5 µm) infrared. Delidding: removing the lid to expose the die. Cards: aifoundry2, aifoundry3 and aifoundry1 card 1 run at 600 MHz; aifoundry1 card 0 idles at 300 MHz and 399 mV on the die. The minion clock: the one clock the 32 compute shires and the master shire run on. The floor voltages: the lowest the firmware accepts, 400 mV on the minion rail and 660 mV on the SRAM rail. A look: the seconds a bare card leaves between the low clock being set and the measurement plan's 75 °C stop. Lock-in: switching a pattern of work on and off at a fixed rate and averaging the images in step with it, which removes everything that does not follow the switching. θ85: the largest thermal resistance at which the die still settles at or below 85 °C. Numbers carry their kind where it matters: measured on our cards, from source read from a vendor document or an outside source, derived arithmetic on those, fitted, model (5th–95th percentile of the Monte Carlo unless stated), assumed, estimate a rough cost or time, inference. References in brackets link to the lists at the end: [R1] and so on to the project's own record and the vendor's documents, [1] and so on to outside sources.
1. The package and the card
Answer. The chip sits in a lidded 45 mm package whose datasheet gives no temperature limit and no thermal resistance. The card takes up to 88 W at its card-edge 12 V input (an auxiliary 12 V header joins after the current sense), has a fan header, and has no power switch the host can operate.
- The package. A 45.0 × 45.0 mm flip-chip ball-grid array (FCBGA) with a full-coverage 44.8 mm lid, 2,494 balls and a maximum height of 3.95 mm (datasheet Fig. 9-1, p. 33) [R1] from source. The lid's top is 1.91 mm above the substrate, which leaves a lid plate about 1.0–1.3 mm thick derived. The dev-card manual's photo shows the bare metal lid (p. 2) [R2]. The die is 25.6 × 22.2 mm [R10]. Esperanto's Hot Chips 33 slides give the same 45 × 45 mm and 2,494 balls, with over 30,000 bumps to the die [3].
- No thermal ratings. The datasheet's thermal section (§9.1) and its ratings (§8.1–8.2) are placeholders: there is no maximum junction temperature and no θJA or θJC [R1]. None was found in Esperanto's public papers and pages either: they give power figures and air cooling, but no junction limit, throttling or thermal trip [3, 4, 5]. The 90 °C stop used in this project is the owner's rule, not a rating.
- Power. The card's power tree shows “+12V, 7.3A Max, 88W Max” at the card edge, the ATX auxiliary 12 V joining after the current sense, a “12V Fan” branch and a thermal diode (p. 5); the assembly drawing shows a P1 “FAN” header and the P7 auxiliary header (p. 3); an LTC4218 hot-swap switch sits on the 12 V input [R2]. The highest draw on record is 87.8 W [R3] measured.
- No power control from the host. None is documented. The card is powered whenever its host is, a reset does not remove its power, and the lab admin power-cycles a host on request. Shutting the host down removes the card's 12 V, but nothing switches the card alone off, and no remote power-on is recorded [R11, R12].
- No protection seen at 600 MHz. On firmware builds 0.20.0 and 0.18.0 nothing has been seen to limit the die once the clock is at 600 MHz. The PMIC's 75 °C / 75 W alarm watches the PMIC's own temperature reading, which reads 0, and on aifoundry2's 0.20.0 build it did not act at 86.9 W (that build's safe state would not change the clock anyway); card 1's 0.18.0 build has a real 300 MHz safe state on the same alarm, read from source and untested. There is no shutdown path [R3, R4]. The stop near 120 °C that the lab lead described is not on record (the effect of overheating, §4.4).
- Not recorded: which heatsink the lab cards carry, whether each has a fan, and whether the sink also covers the DRAM (open question 4 in the full thermal-camera plan) [R13].
2. Why the chip cannot hold a temperature bare
Answer. The heatsink is what keeps the chip's leakage from feeding on itself. A card's idle power rises with its temperature; with the heatsink, the extra heat of each degree is carried away faster than it is made, and the die settles. Without it the thermal resistance is about 3–8 times as large in still air and 1.4–5 times with a fan; the loop gain passes 1 before any balance point, and in the model only 1 of 3,600 draws at 600 MHz settles at idle.
2.1 The thermal resistance, with and without the heatsink
With the heatsink. The six-stage model fitted to aifoundry2 totals 1.47 °C/W [R5] fitted. From each card's idle point, the effective resistance is derived. Card 0's cooling is about 1.3–2 times worse than the others'; on 25 September it read 115–117 °C at 600 MHz with nothing running (66–71 W), until its firmware dropped it to 300 MHz [R3].
The fitted chain, stage by stage
Resistances and time constants from docs/findings/11-thermal-model.md lines 29–31, fitted on
aifoundry2 only; the heat capacity of each stage is C = τ/R. The first stage's 14 J/K matches the package's own heat
capacity (section 3), so it is most likely the package into the paste, the second the heatsink's base, and the third its
body inference.
Without the heatsink. Heat leaves by two parallel paths: from the lid's top into the air, and down through the balls into the board, which then acts as a fin model:
with a lid area Alid of 20.1 cm², θJB of 1–3 °C/W assumed and the board as an annular fin. In still air the lid-top path alone is °C/W and carries only of the heat. The two paths together give θJA of °C/W in still air and °C/W with a 1–2.5 m/s fan model. The vendors' tables for lidded packages with no heatsink sit at and above that: 5.4–7.9 °C/W in still air for AMD's 40–47.5 mm packages, 5.7–7.4 for its four lidded 45 mm ones, and 2.7–4.8 °C/W at 1.3–2.5 m/s [1]; 7.75, 5.80 and 4.98 °C/W for Microchip's MPF500T in its 35 mm package in still air, at 1.0 and at 2.5 m/s [2]. AMD gives its values “for device/package comparison purposes only” [1]. So §2's steady-state analysis carries wider ranges forward that hold both the model and the tables: 4–12 °C/W in still air and 2–7 °C/W with a fan, against 1.47 °C/W with the heatsink.
| Package, lidded, no heatsink | Size | Still air | 1.0–1.3 m/s | 2.5 m/s | 3.8 m/s |
|---|---|---|---|---|---|
| AMD FFVA1517 | 40 mm | 7.9 | 4.8 | 4.1 | 3.8 |
| AMD FLVD1924 | 45 mm | 7.0 | 4.2 | 3.5 | 3.3 |
| AMD FLGF1924 | 45 mm | 5.7 | 3.5 | 2.9 | 2.8 |
| AMD FLGA2104 | 47.5 mm | 5.4 | 3.3 | 2.7 | 2.6 |
| Microchip MPF500T-FCG1152 | 35 mm | 7.75 | 5.80 | 4.98 | — |
| ET-SoC-1 on its card, this page's model | 45 mm | — | |||
| Carried forward in §2–3 | 45 mm | 4–12 | 2–7 | — | |
θJA in °C/W from source. AMD's columns are its 250, 500 and 750 LFM (1.27, 2.54 and 3.81 m/s) on a JEDEC four-layer test board, and its table notes that “All θJA-Effective values assume no heat sink” [1]; Microchip's are still air, 1.0 and 2.5 m/s [2]. The lid helps even without a heatsink: Microchip's MPF300T is 9.50 °C/W in its 29 mm lidded package and 11.21 °C/W in the same package as a bare die [2]. The model's row gives its 5th–95th percentiles, with the fan at 1–2.5 m/s model.
2.2 Leakage feedback, and the runaway condition
Each card's idle board power follows a law fitted to its cooling runs in September: a fixed part plus a leakage part that grows exponentially with the die's temperature [R9] fitted. A temperature T* where the heat made equals the heat removed, T* = Tamb + θ·P(T*), is stable only if a small rise in temperature adds less heat than it removes:
This is the stability test used for power devices: a system is “thermally unstable in case the power generation … rises faster than the power dissipation … over temperature” [10], written dPtot/dTj < 1/Rthj-a for Schottky diodes [11]. In chips, a higher junction temperature “causes further increase on the standby leakage current”, leading to “possibly the thermal runaway” under burn-in [12], and a stable point “may be greater than the operating limit” [13]. An Athlon 64 drew 5.3% more power per 13 °C [14]; aifoundry2's idle board power rises 26% from 70 to 83 °C [R9] derived.
With the heatsink, aifoundry2's loop gain is 0.95 at 80 °C and reaches 1 at 81.9 °C; at the fitted intercept of 22.8 °C the model rests at 62.1 °C and runs away from 98.7 °C, and at 28.0 °C it has no balance point at all [R9]. That last result is a failure of the fit, not a measure of the margin: the fit puts all of the board's power on the die, and every cooling run on record settles (the effect of overheating, §5.2). On the SoC's share of the power, this page's model has aifoundry2 needing θ at or below 1.63–2.17 °C/W for a stable idle temperature model, and its heatsink gives an effective 1.42–1.61 °C/W on board power derived: a margin, but a modest one. Bare, the loop gain at 60 °C is model. Without the feedback, the idle power alone would hold the die above the room.
Lines: how fast each card's idle board power rises with temperature, the derivative of its fitted law (solid over the temperatures it was fitted on, dashed where extrapolated; card 0's law was measured only at 60–64 °C at 300 MHz, so its shape above that is the model's assumption). Bands and line: 1/θ for each kind of cooling. Where a card's line lies above a band, a temperature there cannot settle with that cooling. Lying in or below a band is not enough: a stable temperature must also be a balance point, and bare in still air the idle power alone would hold the die of a 600 MHz card above the room (the paragraph above), so where a line dips into the still-air band at lower temperatures, below about 50 °C, a bare card would still not settle there model. The slopes are of board power; 80–95% of each slope is on the die, which lowers each line by at most a fifth assumed.
The same condition gives, for each card, the largest θ at which any stable idle temperature exists, from the most favourable case (a 22 °C room, the least of the power on the die) to the least (30 °C, the most) model. It sits between the heatsink's value and the bare ranges:
Per card: the largest θ that still allows a stable idle temperature (bar, on the SoC's power), and the effective θ with its own heatsink at its measured idle point (mark, on board power). Shaded: the bare package, 4–12 °C/W in still air and 2–7 °C/W with a fan. Log scale. Card 0 is at 300 MHz and 399 mV; at 600 MHz its bar would lie with the others inference.
The Monte Carlo. Drawing θ across the bare ranges, the room at 22–30 °C and the die's share of the power across its range, At 116 °C the idle laws give 77, 88 and 120 W on aifoundry2, aifoundry3 and card 1 [R4], extrapolated: a bare card at idle would approach its 88 W input at about 105–122 °C. No limit has been seen to act before then on these cards. One is on paper: card 1's 0.18.0 build has a real 300 MHz safe state on the PMIC's 75 W alarm (from source, untested), which by its idle law a bare card 1 would reach near 99 °C derived; the model leaves it out. The hot-swap switch's current limit is not documented.
3. How fast it heats
Answer. In seconds to minutes. The package has a heat capacity of about 12–17 J/K; the heatsink's base and body behind it add hundreds of J/K more. Bare in still air, aifoundry2 reaches 90 °C 1–3 minutes after a cold power-on at idle (card 1 in 41–99 s; aifoundry2 with a fan in 1.3–8 minutes, and 4% of its draws never by 900 s), and in about half a minute under load if the load could start at power-on. A heatsink lifted from a running card at 73 °C would leave at idle and under load model.
The die (0.4–0.7 J/K), the lid (7.8–9.7 J/K, taken as copper assumed) and the part of the substrate that follows them within seconds (6.3–7.8 J/K in all) make a package of J/K model, which matches the 14 J/K of the fitted chain's 1.5 s stage. Without the heatsink's 80 and 256 J/K stages behind it, the idle power alone heats the die at °C/s at first, and °C/s with 23 W more on the die. The model is two nodes, the package and the board around it (30–80 J/K), each losing heat to the air.
aifoundry2 from a cold power-on.
The time to a given temperature, from a cold card at room temperature, for each card and cooling model:
3.1 The host-boot window
A lab card is powered whenever its host is, and nothing can run on it until the host has booted and loaded the driver. How long that takes is not recorded; this page assumes 30–90 s from power-on to the first telemetry reading assumed. By then the bare die, in still air, is already at on aifoundry2 and on card 1. From there to 90 °C: So in still air bare aifoundry2 would allow one load of about per power cycle, and card 1 about model, boot assumed. The die keeps heating throughout, so only a short modulated run fits: lock-in of a pattern against its complement through a taped lid needs about 10 s at 600 MHz (section 5.4). Short of shutting the whole host down nothing switches the card off, and no remote power-on is recorded either.
The die at the end of the host boot, and the seconds left to 90 °C, per card
Each row: 300 draws; the card idles through a boot of 30–90 s, then the load starts. “Already past 90”: draws in which the die was at 90 °C or more when the host came up (150 means it had passed 150 °C, where the model stops).
4. What a sensor would see
Answer. With the heatsink off: a plated lid at nearly one temperature, which a camera reads mostly as the room's reflection unless it is taped. With it on: nothing of the chip, since the heatsink covers the lid. The only view that resolves shires is of a delidded die under an infrared-transparent cooler, which needs a card that may be lost and, as in the published setups, a mid-wave camera; in the long-wave band of the lab's camera it needs a germanium, zinc-selenide or similar window and a coolant that transmits there, for which a thin mineral-oil film is the untested candidate (section 4.2). A thermocouple on the lid adds one accurate point, not a picture. All of this is about a single frame: a pattern of work switched against its complement and read by lock-in does show through a taped lid, coarse patterns within seconds at 600 MHz model (section 5.4).
How far a hot spot spreads sideways (decay length, 5th–95th percentile, log scale), against the 3.7 mm spacing of the shires and the die's 22–26 mm width model. A view resolves single shires only if its length is near or below the shire spacing.
- The bare lid. As a plate cooled by still air (copper assumed), the lid alone would spread heat over mm, many times its own 45 mm width, so the full thermal-camera plan's estimate of about 200 mm (its target 9) holds [R13]. With the heatsink off the heat is forced down into the substrate, and the die and lid together still spread it over mm, about the die's width. A 1 W load in one shire would leave a broad bump of °C over the lid at 2 mm and °C at 10 mm, so “no spatial structure at all” slightly overstates it, but in the raw image no shire could be told from its neighbour. Sadiqbatcha et al. argue that a thin spreader “will distribute the heat across its surface” but “the spacial locations of the underlying heat-sources do not change”, and recover them in simulation with a 2D Laplacian [15]. Here the lid is 1.0–1.3 mm thick and that bump is about a degree, so whether the method would work on this lid depends on the camera's noise, and on taping the lid inference.
- Reading it. The model takes the plated lid's emissivity as 0.05–0.3 assumed. In KIT's emissivity test of a packaged chip, the metal read “very inaccurate (emissivity ~0.01). Most heat measured is reflection from surroundings” [8]. An infrared camera therefore reads bare metal low and mostly sees the room, unless the metal is covered with tape or paint; KIT covered it with masking tape “with an emissivity of 0.92”, and calibrated the camera against the chip's own thermal diode [8]. An emissivity below the model's range would only raise the bare θ a little, through less radiation inference. A thermocouple reads the lid correctly, but one lid temperature adds little to the die's own sensor mean, which the host already has.
- With the heatsink on. The lid's decay length falls to mm (the full plan's target 9 estimates about 3.7 mm), short enough to show structure, but the heatsink covers it.
- A delidded die under an IR-window cooler. Liquid flowing across the bare die under a window transparent to infrared shortens the length to mm, or mm with 5–10 Hz lock-in excitation: that resolves single shires. Every published cooler of this kind that states its band was imaged by a mid-wave camera or microscope, at 3–5, 3.4–5.1 or 2.5–5.1 µm, with mineral oil or a fluorinated liquid as the coolant, and sapphire where a window is named (section 7.4). The lab's planned camera, the Thermal Master P3, sees long-wave infrared, 8–14 µm, at 25 Hz [R13], where sapphire and fused silica fail, Czochralski silicon transmits poorly and fluorinated coolants likely absorb; it would need one of the windows and a coolant of section 4.2, a combination none of the setups in section 7.4 used. Realistically this is a job for a mid-wave camera.
- Long-wave without a window. KIT opens the package and cools the bare silicon from behind, “through the PCB”, with a Peltier element [8], imaging the exposed silicon from above with an 8–14 µm camera [16]; a UC Riverside group does the same with a FLIR A325sc (7.5–13 µm) on laptop processors, an RTX 4060 and a Coral TPU, and notes that through-PCB cooling's “efficiency is significantly reduced” [15]. On this card the board path starts with θJB, assumed at 1–3 °C/W, already near the °C/W at which an idle card can still settle, so rear cooling may not hold it inference.
4.1 A thermocouple on the lid
Intel defines case temperature at the geometric centre of the lid and solders a 36-gauge type T thermocouple into a groove cut in the lid there [9], and AMD warns against any thermocouple between the package and the heatsink [1]. A calibrated 36–40 AWG type T bead at the lid's centre has an absolute error of °C (±2–3 °C for an uncalibrated type K), resolves steps of 0.02–0.1 °C on a 24-bit logger and responds in 0.1–0.5 s model. The model's budget is for a bead in a groove in the heatsink base, pressed on the lid, where part of the paste's temperature drop can fall between bead and lid; soldered into the lid, that term goes, so the estimate is if anything high inference. The die's mean sits above the lid by θJC·P, about °C at idle and °C under the +23 W random-data load, with θJC taken as 0.03–0.10 °C/W assumed. The host sees only a whole-degree mean of the 34 shire sensors and an anonymous peak-hold [R3, R4].
What it adds: a physical point between the die and the cooler, which separates the chain's first two stages from the heatsink's; θJC to ±30–60% under load; an absolute check on the sensors' mean; and a log that runs when the host does not. What it does not add: anything spatial. Re-pasting the heatsink resets that card's fitted thermal constants, so the groove belongs on a card whose heatsink is coming off anyway. Intel solders the bead in at 150 ± 3 °C [9]; done on a card, the package would see that temperature, above the 115–117 °C card 0 read on 25 September inference, which is one more reason the lid groove needs the lab's consent.
4.2 Windows and coolants for the lab's long-wave camera
A camera looking at a cooled die sees it through the window and through the coolant film under it, and both must transmit in the camera's band. For the P3's 8–14 µm [R13] from source:
| Window | Transmits | At 8–14 µm | Price, 50 mm |
|---|---|---|---|
| Germanium | “the whole of the 8-14 micron thermal band”; uncoated it loses 53% to reflection; it “becomes opaque at all wavelengths a little above 350K” (77 °C) [17] | yes, anti-reflection coated and kept below about 77 °C | about $919, 1 mm, coated for 8–12 µm [18] |
| Zinc selenide | 0.5–20 µm; “slightly toxic” [19] | yes | about $875, 2 mm, coated for 8–12 µm [20] |
| Zinc sulphide (FLIR grade) | 1.0–13 µm [21] | yes, to 13 µm | — |
| Chalcogenide glass (AMTIR-1) | 0.75–14.0 µm [22] | yes | — |
| CVD diamond | “a compelling choice for some more extreme far infrared (8–14 μm) window applications” [23] | yes | — |
| Barium fluoride | to about 12.5 µm [19]; “useful” at 0.265–10 µm [24] | partly | $264 at 50.8 mm [24] |
| Calcium fluoride | cuts off at 8.7–10.5 µm, by thickness [19] | partly, the short end | about $295, uncoated [25] |
| Sapphire | 0.17–5.5 µm [26] | no | — |
| Fused silica | poor or unusable in the long-wave band [19] | no | — |
| Silicon (Czochralski) | used “primarily in the 3 to 5 micron band”, with an oxygen absorption band at 9 µm [27]; lattice absorption about 1 cm⁻¹ at 9 µm and over 2 cm⁻¹ at 11–16 µm [28]; transmission depends on doping [29] | poorly | — |
| KBr, NaCl | transmit, but “soluble in water” [19] | yes, but water-soluble | — |
The Edmund prices are from search-result listings in September 2026 and were not checked on Edmund's pages, which block automated fetching; the EKSMA price is from its page. Prices are for the window alone.
Coolants. Fluorolube, a fluorocarbon mulling agent, has “strong carbon-to-fluorine bond absorptions from 1300 cm−1 onwards to 400 cm−1” [30], that is from 7.7 to 25 µm; other perfluorinated coolants likely absorb across much of 8–14 µm too inference. They suit the mid-wave better: FC-70 transmits over 90% there, though at the tested film thickness it too appeared opaque to the Stanford group's mid-wave microscope [31]. Mineral oil (Nujol) has its major absorption peaks at 2950–2800, 1465–1450 and 1380–1300 cm⁻¹, that is 3.4–3.6, 6.8–6.9 and 7.2–7.7 µm [32], and serves spectroscopy from 1370 cm⁻¹ into the far infrared, except for one band at 720 cm⁻¹ (13.9 µm), at the long edge of the P3's band [33]. A thin mineral-oil film is therefore the coolant to try in the P3's band, but it is untested: a spectroscopy mull is a paste pressed between salt plates [30], not a flowing layer thick enough to carry the heat, and how much such a layer absorbs across 8–14 µm is not known inference. Water is “not transparent” [14]. The film must also be thin for accuracy: with Galden HT-170 the error is about 0.1 °C under a plenum of less than 500 µm and up to 43 °C for a 2 mm channel, where the film is “effectively opaque” [31].
5. Slower clocks: how low can it go, and does it help?
Answer. The lowest minion clock the firmware can set is 100 MHz, six times slower than 600 MHz. Nothing lower can be set without new firmware, and 10 MHz would save only more than 100 MHz from source; model. A slower clock barely changes what a bare card has to shed: most of its idle power is leakage, which the voltage and the die's temperature set, and power that no minion clock drives. On aifoundry2 at a 60 °C die, 600 → 100 MHz with the voltages left as they are removes of ; the firmware's lowest voltages and a slower NoC bring the card down to model. Even there, no clock holds a bare card: settling below 85 °C needs a thermal resistance of or less. In still air no draw settles at any clock; with a fan on the bare lid of draws do on aifoundry2 and aifoundry3, a share set mostly by the assumed range of the fan's cooling; with a 3–6 m/s blower the measurement plan's model brings of cold power-ons through its gate, which asks for 80% model. A slow clock buys little time, since the card idles at 600 MHz until the host is up and the low point is set, and it costs signal: the heat that marks where the work runs falls with the clock. Through a taped lid at 100 MHz, lock-in shows a single shire or a coarse pattern within about 10 s and a checkerboard in about , inside the look a fan leaves in the better power-ons model; such a look is possible only behind an automatic power cut and after the measurement plan of section 5.6. The views that hold a temperature are with the heatsink on, the chip's own sensors read with lock-in (one usable sensor on the stock firmware; a per-shire map needs a rebuilt boot loader), or a delidded die under a cooler.
The owner's question (30 September 2026, about 15:45 PDT), verbatim
“For the heat sink issue, can you double-check that maybe I can run it much lower frequency, like how low can I go? Can I go at 100 MHz? So, um, yeah, find some other alternatives. I want to consider maybe doing something extreme, like 10 MHz, and get the envelope of what's possible. Ideally, I would run at the slower speed, maybe 10 times slower, but I would be able to run it without the heat sink and point the thermal camera and [see the] arrange[ment] of the computation.”
5.1 What the firmware can set
- One clock drives all the minions. Every firmware build boots the minion shires, the 32 compute shires and the
master shire, on one “step clock”, a PLL in the service processor, at 600 MHz, and turns the per-shire PLLs off; the
shire caches and the master minion run on the same clock, and which clock the spare shire gets is not established
[R15] from source. The management command
DM_CMD_SET_FREQUENCYreprograms that PLL and changes no voltage from source. It accepts only a frequency that is an exact entry of the PLL's mode table, and below 300 MHz the table has 100, 125, 150, 166, 175, 200, 225, 250 and 275 MHz; 100 MHz is mode 62, a 1,100 MHz oscillator divided by 11 from source, current tables inference for the lab's builds. aifoundry3's boot service already sends it (-f 600,400) at every boot [R15] measured. - Nothing below 100 MHz. No table entry, management command or boot-loader path reaches a lower clock (the
minion-debug commands' memory writes reach only DRAM, a kernel stack and trace buffers), and the two other inputs a
shire can select, the PLL's bypass and the shire mux's reference position, are 100 MHz as well
[R15] from source. The hardware's divider
field would allow it: 10 MHz needs a post-divider of 107 at the tables' lowest oscillator frequency and up to 400 at
their highest, against a 9-bit field derived,
pll_modes.py. But only a new signed boot loader, which the lab's cards have never run, or a debugger could program one. Esperanto's slides give the minions an operating range of “300 MHz TO 2 GHz” [3], and the PLL vendor's generic brief gives an output of “4MHz - 6GHz” and a “Divider Range 1-256” [44]: the floor is this chip's tables, not the PLL's design inference. - The voltages.
DM_CMD_SET_MODULE_VOLTAGEsets each rail on its own. Builds 0.20.0 and later accept a minion rail of 400–620 mV, an SRAM rail of 660–850 mV and a NoC rail of 400–600 mV; 0.18.0 (card 1) checks nothing and writes every NoC setting to flash [R15] from source. The firmware's “safe state” names 300 MHz at 400 mV on the minions and 660 mV on the SRAM, constants marked “TODO: to be removed and replaced with vmin lut value” (650 mV on the SRAM in 0.20.0, below that build's own 660 mV check): a lookup, not a validated operating point [R15] from source. Card 0 idles near it, at 300 MHz and 399 mV on the die: 17.97 W at a 62 °C die (card0_guard.py) measured. aifoundry2 runs 600 MHz at 517–519 mV on the die (517 with a kernel running, 518–519 idle) [R7], and aifoundry2's and aifoundry3's SRAM rails are set 40–45 mV above that rail's floor measured, from source. Esperanto puts the minions' “sweet spot” “between 300 and 500 mV”, a modelled figure [4]; no published SRAM floor was found, and SRAM's energy-optimal voltage sits about 100 mV above logic's [45]. The regulator is not the limit: the module that feeds the minion and NoC rails [R15] is specified down to 0.25 V [46]. - Which card. Only the firmware's governor ties voltage to clock, and only at the points of each card's operating-point table (aifoundry2 has been seen at 600, 700 and 800 MHz [R7]; no card's full table has been read). On aifoundry2 the governor would put 600 MHz back at its next idle or when its thermal loop exits; aifoundry3's governor is latched, so a clock set there stays, barring the race below (the plan starts from a freshly booted card); card 1's voltages must never be set; card 0 is excluded. So aifoundry3 is the card for any trial [R15, R16] from source, inference.
- Not established. No lab card has run below 300 MHz, beyond the microseconds each PLL change spends bypassed at 100 MHz. At 100 MHz the master minion's handshakes with the compute minions take six times longer, while its timeouts for them, 5–10 ticks of a wall-clock timer, do not stretch: about 52–105 ms if a tick is the 10.5 ms that the measured 1.05 s heartbeat implies inference. Reading the source also turned up a race: the clock command and the governor's loop reprogram the same PLL with no lock between them, and the command's task can interrupt the governor's anywhere, so a clock set while a governor loop runs, even during the loop's own voltage change, can be overwritten by the governor's stale target. The current source still has it [R16] from source, each step inference: the window can be hit never seen on a card.
The rest of this section compares three voltage policies assumed. Voltages left: the clock command alone, so the rails stay at the card's 600 MHz values (518 mV minion and 704 mV SRAM on aifoundry2). Floor: minion 400 mV and SRAM 660 mV at and below 300 MHz, the safe state's pair (398 and 660 mV on the die in the model); at 400 MHz a straight line to the card's 600 MHz point gives 438 mV and 675 mV, and what is safe there is not established. Floor + NoC: the floor, with the NoC also at 200 MHz and 400 mV (398 mV on the die in the model), values the firmware accepts but the vendor's operating-point validator stops short of (485 mV), so untested [R15].
5.2 What a slow card draws, and the floor no clock removes
The envelope model (lowclock_calc.py) splits each card's idle power, rail by rail, into what the minion
clock drives, leakage, and what neither touches. A fit to one card's cooling runs leaves the clock's share wide open, so
the model shares one minion clock tree across the three 600 MHz cards (same design, scaled by V²), and it fixes how
leakage falls with voltage from two anchors in the record. aifoundry2 idles at 35.0 W at 800 MHz (a 64–65 °C die)
against 28.1 W at 600 MHz (63–68 °C), compared through its idle law at 64.5 °C, with its SRAM rail stepping from 704 to
830 mV along with the clock; and card 0's minion rail draws 4.09 W at 300 MHz and 0.399 V against aifoundry2's 8.30 W
at 600 MHz and 0.519 V, both at 62 °C, on chips whose NoC rails agree within 2% measured. Only
of the prior draws pass both anchors and the rail fits, and the anchors pull against each
other: card 0's firmware state may not be a pure change of clock and voltage model, inference. The fit gives a minion clock
tree of on the die at 600 MHz, minion leakage at 398 mV of its
value at 518 mV, and SRAM leakage at 660 mV of its value at 704 mV
fitted; transistor physics alone, from N7's drain-induced barrier lowering, “DIBL is ~40 mV/V”
[47], would give 0.65 derived. The clock tree's
size rests on the exponential form of the leakage law in temperature: Arrhenius forms fit the same cooling runs and leave
up to about 3.5 W of the minion rail that does not follow temperature [R16]
fitted, which Stage A of the measurement plan would settle. A clock cuts only the dynamic term,
“P = Cf V² + Pstatic” [48]. At a 60 °C die aifoundry2's idle
model.
- 100 MHz with the voltages left removes only the idle clock tree, , and leaves the rise of idle power with temperature where it was, model.
- The floor voltages remove more, and the NoC at 200 MHz and 400 mV another . That leaves , of which about is on the die, and the die's power rises by model.
- 10 MHz, if it could be set, would remove more at the board: nine tenths of the of minion and SRAM clock tree left on the die at 100 MHz and the floor voltages, with its regulators' loss model.
- The floor no clock removes, with the clock taken to zero: model.
- Under load the clock matters more. A random-data pattern on 16 shires adds model: the load stops mattering, the idle floor does not.
Against Esperanto's own measurement, and Intel's 10 MHz chip
The only published breakdown found is Esperanto's: an idle card at 20.0 W, with 5.8 W on the minions, 1.6 W on the NoC and 1.7 W on the SRAM, clock and temperature not given [49]. aifoundry2's model has a 5.8 W minion rail at a 46 °C die, where its NoC rail is 1.75 W, its SRAM rail 0.91 W against Esperanto's 1.7, and the card 21.5 W model. 46 °C lies below every measured point of aifoundry2 (62–84 °C), so these are its fitted laws extrapolated. Intel's Claremont ran from “1.1V/741MHz/445mW to 380mV/10MHz/1.5mW”, its leakage growing from 3% to 50% of its power on the way: it reached 10 MHz by lowering the voltage with the clock [50].
5.3 Does a slow card hold bare?
Not in still air or with an ordinary fan, at any clock; a strong blower is marginal. The lower panel of the chart above gives θ85, the largest thermal resistance at which the die settles at or below 85 °C, on the SoC's own power. , while the bare package offers 4–12 °C/W in still air and 2–7 °C/W with a fan model.
- Still air: no draw settles below 85 °C, at any clock, settable or not, on any card, under any voltage policy model.
- A fan on the bare lid: the best settable point, 100 MHz at the floors with the NoC lowered, settles in . No clock, settable or not, reaches 90% of draws model.
- A blower of 3–6 m/s assumed would bring the package to , the board path setting the floor. At aifoundry3's lowest settable point the measurement plan's model then gives [R16] model.
- With the heatsink every clock settles at idle. Under a sustained random-data pattern on 16 shires model.
- An IR-window cooler on a delidded die holds every clock, at about 27–30 °C model.
What a slow clock buys: a little time, and only with a fan. A bare card is powered whenever its host is, so it idles at 600 MHz through the host's boot (30–90 s, as in section 3.1), and setting the low point (power management off, the voltages, the clock) takes another 20–60 s at that idle assumed. At the end of the boot the die is already at . The look that is left, from the low point being set to the measurement plan's software stop at a 75 °C die, with the pattern running on aifoundry2 model:
5.4 What the camera would see
The imaging model (imaging_calc.py) is a layered conduction model of the die, its paste, the lid and the
heatsink over the substrate and board, with the work spread evenly over each active shire's 3.7 mm tile, and four
patterns of fp32 random-data work: a checkerboard of 16 shires against its complement, one shire switched on and off, a
2×2 block moving between two corners, and the west half against the east half. The camera is the lab's Thermal Master
P3 (noise-equivalent temperature 24–35 mK, 25 frames a second) at 60 mm, where a shire spans about 22 pixels
[R13]. A pattern counts as seen at three times the noise of a pair of
shire-sized regions model. Only the switching power per shire sets the signal:
model. The view matters more than the clock:
The contrast of each pattern in a difference image of two settled states, in mK, at 600 / 100 / 10 MHz, the voltage at its floor below 600 MHz (aifoundry2's voltages) model. For the camera it is what the camera reads, the temperature contrast times the surface's emissivity, taken as 0.90–0.95 for black paint, 0.74–0.92 for polyimide or masking tape on the lid and 0.01–0.1 for the bare plating assumed [R13, 8], and through an IR-window cooler also the window's and the oil film's transmission, assumed. For the sensors it is the true contrast at the transistors (10 MHz not computed). model.
- A taped lid, heatsink off. The lid (copper, assumed) spreads the heat: model. model. Bare plating is 10–100 times worse, mostly reflecting the room [8]: tape it.
- A delidded die, coated black: model. Bare silicon, whose emissivity is uncertain (0.1–0.7 assumed), needs . An IR-window cooler adds no contrast; it keeps the die alive, and the camera then reads the die through the window and the coolant film model, transmission assumed.
- The chip's own 34 sensors, heatsink on, see the same contrast between neighbouring shires as with the heatsink off, about 1.4 K per watt per shire for a checkerboard, because the lid does the spreading, not the heatsink model. The one spatial measurement on record, a 4-shire block beside the I/O shire against the far corner with the heatsink on, read 0.67 °C on aifoundry3 (99% interval 0.37–0.97) and 0.97 °C on card 1 (0.52–1.41), in whole-degree readings measured (the heat-placement study), against the model's steady : the intervals take in the model's values, and its contrasts are, if anything, low. Read shire by shire, as raw codes, the sensors would show every pattern within 10 s at every settable clock, and down to about 7–10 MHz with 10 minutes of lock-in model; raw-code noise assumed, provided the codes' own noise (0.02–0.09 °C assumed) dithers their 0.061 °C steps; at the low end of that range the amplitude read varies by about ±20% with the starting temperature (10th–90th percentile) model. But the stock firmware gives the host only a whole-degree mean, an anonymous peak and the I/O shire's own sensor [R3], and a per-shire readout needs a rebuilt boot loader. The I/O shire's sensor alone sees a 2×2 block beside it, and its whole-degree readings average below a degree only while the die's temperature crosses degree steps: model. Thermal maps have been almost fully reconstructed from a minimal number of on-die sensors [51]; this chip has one per shire.
- Lock-in. Switch a pattern against its complement at 0.25–1 Hz. For the checkerboard and the moving 2×2 block the total power stays constant, so the mean temperature, the governor and the leakage stay put, and the contrast doubles; the half die (18 shires against 14) and one shire switched on and off also move the mean. The thermal diffusion length in silicon is about 16 mm at 0.1 Hz and about 5 mm at 1 Hz model: slow modulation keeps the signal, while the 5–10 Hz that sections 4 and 7.4 assume sharpens the image at the cost of signal. On a bare card the die keeps heating through the look (section 3), so the states must be taken in an A-B-B-A order and the slow rise removed with a polynomial fit before the difference; the times here assume a steady scene with a 4-fold allowance for drift inference. Lock-in detects heat sources “of a few µW corresponding to a local temperature modulation of a few µK”, and “permanently existing heat sources in the device, which are not affected by the trigger signal, do not appear in the lock-in thermogram” [29]; activating firmware functions periodically has located hardware blocks “at the die level on a modern SoC” [52], and a cooled camera reached 35 µK after 16 minutes [53]. Any run longer than 10 s needs the owner's approval: the lab caps a device hold at 10 s.
5.5 The alternatives, ranked
For the owner's aim, seeing where the computation runs, from the safest to the most drastic:
| # | Alternative | What it would show model | Holds a temperature? | Risk | Needs |
|---|---|---|---|---|---|
| 1 | Keep the heatsink; use the chip's 34 sensors as a coarse camera, with lock-in of a pattern against its complement | Stock firmware: a 2×2 block beside the I/O shire, reliably only with a slow dither of the chip's power over whole degrees, in . With each shire's raw reading: every pattern within 10 s at every settable clock, at one pixel per shire | yes for the stock 2×2 pattern and for runs of 10 s or less; a sustained 16-shire pattern only at (5.3) | none on the stock path; a rebuilt boot loader has never run on the cards | the owner's approval for runs over 10 s; for the per-shire map, a boot loader built and signed to report it (the same image could carry clocks below 100 MHz, which would save only about 0.2 W) |
| 2 | Keep the heatsink; the camera on the back of the board, with lock-in (the thermal-camera plan's view) | coarse patterns, blurred by the substrate and board; not modelled here inference | yes | none | the camera on site |
| 3 | Short looks at a bare, taped lid at 100 MHz and the floor voltages, after the measurement plan, behind a latching mains cut at 80 °C, with forced air and someone on site | one shire, a half die or a moving 2×2 block within 10 s, a checkerboard in about , at a small fraction of a bare die's shire-to-shire contrast | no; with a fan, ; with a blower, of power-ons pass the plan's gate, which asks for 80% | moderate: the card, and a host that may hang if the cut is on its 12 V; taking aifoundry3's heatsink off resets its fitted thermal constants and takes the demo card out of service | Stages A and B of the plan; the owner, the lab lead and the gates of section 5.6 |
| 4 | A delidded card under an IR-window cooler, at any clock | every pattern within 10 s at 100 MHz and above, steadily, through the window and coolant film; for the lab's 8–14 µm camera, a germanium or zinc-selenide window and a coolant that transmits there (section 4.2) | yes, at about 27–30 °C | the card, if delidding fails | a card that may be lost; weeks; thousands of dollars |
| 5 | A delidded card, coated black, at 100 MHz and the floor voltages, with a fan, in short looks | single shires in a difference image of two 1 s states; every pattern within 10 s | no better than a bare lid inference | high: the card | a card that may be lost, AI Foundry's consent, the power cut of 3 |
| 6 | Below 100 MHz (25 or 10 MHz) | nothing the 100 MHz point does not; it saves about 0.2 W more | no | a firmware image never run on the cards | a new signed boot loader (the same kind of image option 1's per-shire map needs) or a debugger |
Bare looks, settling and lock-in times are the envelope and imaging models' (sections 5.3–5.4); the blower figures are the measurement plan's model [R16]. Option 1's stock path and option 2 fit the lab as it is; 4 and 5 need a card set aside, as in section 6. Options 1 and 6 both hinge on whether the cards accept a rebuilt, signed boot loader, which has never been tried.
5.6 The measurement plan, in brief
The record already predicts most of the answer, but two numbers decide the bare verdict and neither has been measured: how much idle power the clock drives, and how steeply leakage falls with voltage. A plan to measure both, heatsink on, is written out in full in LOWCLOCK-PLAN.md [R16]. It is a design: nothing in it has run, and every stage needs the owner's go-ahead and the lab lead's consent, Stage B and the bare step each separately.
- The card is aifoundry3: its latched governor leaves a set clock alone, its boot service already runs the clock command and restores 600 MHz after any host reboot, it is a single-card host, and its firmware checks voltage ranges and does not write them to flash. It is also the lab's demo card.
The stages and the safety rules
- Stage 0, read-only (about 20 minutes): identity, the card's operating-point table (stop if 100–400 MHz are in it), and whether a kernel has run since boot. The lab lead stops the demo service first.
- Stage A, heatsink on (): power management off, then the minion clock stepped 600 → 400, 300, 200 and 100 MHz at the card's own voltages, in four passes in a Latin-square order, with a checked kernel launch verifying each clock by its work rate. The step 600 → 100 MHz is predicted at model.
- Stage B (; a separate approval): at 100 MHz, the minion rail 525 → 400 → 525 mV and the SRAM rail 700 → 660 → 700 mV, then 10 minutes at 400/660 mV, predicted at at a 56 °C die model. Its first write sets the value already there, to test the write path. Its leakage exponent decides between the plan's model and this page's.
- Safety: a hard-coded list of the only values that may be written; stop rules on the driver's error counters, which any user can read, and on the kernel log; the restore puts the voltage back before the clock; during Stage B someone who can power-cycle aifoundry3 within about 30 minutes, after a link retrain hung another host on 30 September.
- The bare step runs only if the measurements agree with the models and the model, re-run with them, gives at least 80% of power-ons that stay below an 80 °C stop and settle at or below 70 °C. It needs someone on site, a thermocouple on the lid with its own logger, an automatic mains cut at 80 °C that stays off until reset by hand, blower airflow, and a software stop at 75 °C. On today's predictions no cooling class passes: model. Taking aifoundry3's heatsink off and refitting it resets the thermal constants fitted with it, and takes the demo card out of service meanwhile.
- Two firmware hazards turned up in the design, both read from source and never seen on a card: a clock missing from the operating-point table makes a running governor loop log an error on every pass, and the firmware sends each error to the host as an event; and the race of section 5.1. The plan avoids both, and the race is a report for the firmware's maintainers.
6. What would work, ranked
Answer. Keep the heatsink, and tape a thermocouple to its base during a visit already needed, never between it and the lid. For the case temperature as Intel defines it, and only with the lab's consent and at moderate risk, solder one into a groove in the lid. Put emissivity tape wherever the camera looks. Anything that takes the heatsink off for good belongs on a card set aside for it, on a bench, not in a lab host. A slower clock does not change this (section 5). The safe way to see where the work runs is the chip's own sensors, heatsink on, read with lock-in: on the stock firmware that is one sensor beside the I/O shire, and a per-shire map needs a rebuilt boot loader. A short bare look at 100 MHz is possible only with forced air, in some power-ons, behind an automatic power cut and after the measurement plan.
Who does what. AI Foundry's people do all physical work on site; the lab admin powers hosts down and up; the lab lead and AI Foundry decide what happens to the cards. The lab's rules still hold for any run: no resets or configuration changes, the card lock, and at most 10 s holding a device (AGENT.md §5). Tape goes on with the card off, and it is polyimide rather than vinyl (pitfall 10 in the full camera plan) [R13]. If a heatsink comes off, AMD's guidance for its own lidded packages applies: warm a large part to about 40–60 °C and twist the heatsink free; wipe the residue off with a cloth wetted with solvent, as AMD's removal steps do (isopropyl alcohol is on its list), without flooding the package, since AMD's soldering chapter warns that washing solvents “can compromise the lid adhesive”; and refit with thermal paste (a heatsink without it “is not sufficient”) at 20–50 psi with four-corner mounting [1].
Why card 0 for the groove. aifoundry1 card 0 is excluded from the measurement campaign [R14] but works at 300 MHz. It is the lab's card and shares aifoundry1 with card 1 and the CI runners, so any physical work there takes both cards down. The lab-problems page already asks for an on-site check of its fan, heatsink seating and thermal paste (request SH1): its heatsink is likely to come off then anyway, and its cooling is already suspect (1.3–2 times worse than the others'), so re-pasting it loses no fitted constant we rely on. Cutting its lid is a permanent change, and destructive use of any card needs AI Foundry's consent. The full camera plan's line (under E-T20) that “aifoundry1's two are unusable” is out of date [R13].
7. What others have done
Answer. The outside record agrees with the model. Lidded packages of this size are rated at 5.4–7.9 °C/W with no heatsink in still air (section 2.1). In published bare runs a processor that throttled itself kept running, and two burned out within a second, one with no protection and one whose board-level protection reacted too slowly; Intel requires a tripped Atom 330's power removed within half a second. Case temperature is measured with a thermocouple in a groove in the lid, not under the heatsink. Die maps come from opened chips under infrared-transparent coolers imaged in the mid-wave band, or cooled through the board and imaged in the long-wave. Esperanto publishes no temperature limit for this chip.
7.1 Running a processor without its heatsink
- Tom's Hardware, 2001. The site pulled the heatsinks off running processors [6]. An Athlon 1400, a bare die with no thermal protection (Socket A dies were exposed [34]), burned out: “In less than a second Athlon 1400 dies the heat death”, its core reaching, by the article's account, 370 °C. An Athlon MP, whose motherboard read its thermal diode, crashed “a split second after the heat sink had been taken off” and then died near 300 °C; the diode could handle “only 1 degree/s”. A Pentium 4 kept running because “the thermal unit throttles down the clock … until a safe temperature has been reached”, and a Pentium III hung but survived [6].
- The trip, and the power behind it. Intel's Atom 330 stops execution through its thermal trip (THERMTRIP) “when the junction temperature exceeds approximately 125°C”, but “leakage current can be high enough such that the processor cannot be protected … without the removal of power”, so the supply “must be turned off within 500 ms to prevent permanent silicon damage due to thermal runaway” [7]. NVIDIA's GPUs report a slowdown temperature, at which they “begin slowing itself down through HW”, and a shutdown temperature [35]. No rigorous public time-to-trip data for bare GPUs was found.
- A modern bare run. A Ryzen 3 4300U reportedly ran Crysis with no cooler only after its limit was cut “to 90C down from 100C”, and scored 327 in Cinebench R15 [36].
- A model. KIT's model of an Alpha processor runs at 90 °C cooled and at “> 200◦ C (when both packaging and heat sink are removed)” [8].
- This card. At 600 MHz on the lab's firmware nothing has been seen to act (section 1), and no public source gives a limit [3, 4, 5]. The Athlon MP is the closer case, a protection that existed but acted too slowly: here the host reads only a whole-degree mean [R3], and nothing short of a host shutdown cuts the card's power inference. The model's minutes to 90 °C, against the Athlons' second, owe much to the lid and package's 12–17 J/K (section 3) inference.
7.2 Measuring case temperature
- Intel. “The measurement location for TC is the geometric center of the IHS”: the method cuts a groove in the lid, lays an “Omega, 36 gauge, ‘T’ Type” thermocouple in it and solders it at 150 ± 3 °C [9]. This is option 3, which on a lab card needs the lab's consent.
- AMD. “Thermocouples should not be present between the package and the heat sink”: one there “might cause additional mechanical and/or thermal stress” and makes the paste layer “thicker and or uneven” [1]. This is why option 2 tapes the thermocouple to the heatsink's base or fins, outside the joint, and why this page does not put a bead in a groove in the heatsink base, in the joint.
7.3 Taking a heatsink off, and delidding
- Refitting. AMD's guidance for its FPGA packages: a heatsink without thermal interface material “is not sufficient”; the pressure should be “20 to 50 PSI”, “the same for both lidded and lidless”, with four-corner mounting and a bracket; remove it by twisting, warming large parts “to about 40°C–60°C” first (Laird's advice for phase-change material, which AMD quotes), and wipe the residue away with a cloth wetted with toluene, acetone, an isoparaffin or isopropyl alcohol; its soldering chapter warns separately that washing solvents “can compromise the lid adhesive” [1].
- Delidding. The exposed dies of Socket A processors “were susceptible to damage” when coolers were installed incorrectly or systems were handled roughly [34]; soldered lids need 165–180 °C to lift, and delidding “voids the manufacturer's warranty” [37]; “With one slip, you can damage the surface-mounted components” [38]. A bare-die package also runs hotter with no heatsink: Microchip's MPF300T is 11.21 °C/W as a bare die against 9.50 lidded [2].
7.4 Imaging a working die
Published power maps of working processors, and what each took from source:
| Group | Chip, window, coolant | Camera band | Heat removed, accuracy |
|---|---|---|---|
| IBM (Hamann et al., 2007) [39, 40] | “effectively cooled using an IR-transparent heat sink”; the patent's window “polished silicon, quartz, sapphire or diamond”, coolant “perflouro-octane, perflouro-hexane, octane, or hexane” (or “water or a cold gas”) in a 0.1–20 mm duct | not stated in the abstract or patent; quartz and sapphire pass only the mid-wave, and a companion paper reportedly used an InSb (mid-wave) detector | “up to 200 Watts/cm2 with a corresponding temperature increase of 70 degrees C” |
| UC Santa Cruz (Mesa-Martínez et al., 2007 and 2010) [14, 41] | a “non-lidded” Athlon 64 under mineral oil “designed for infrared spectrography”; water rejected; “a 3mm thick sapphire window” in 2010 | FLIR SC-4000 InSb, “3-5μm”; “Si has a fairly uniform 55% transmittance from 1.5μm to 6μm” | “up to 100W” |
| Brown (Reda, Nowroz, Dev) [42, 43] | oil in a “1 mm” channel between “two infrared-transparent sapphire windows”, the oil doubling as the thermal interface | FLIR SC5600, “2.5 – 5.1 µm” | oil at “5 m/s” and “20 Celsius”; “about 90% of the heat flows upward” |
| Stanford, with AMD (Hom et al., 2012) [31] | Galden HT-170 over the die under a sapphire window; FC-70 transmits over 90% in the mid-wave, though at the tested film thickness it too appeared opaque | mid-wave, 3.4–5.1 µm (an infrared microscope) | “~0.1 °C error … if the fluid plenum height is less than 500 μm. For a 2 mm channel, the error can be as high as 43°C” |
| KIT (Amrouch and Henkel, 2015) [8, 16] | an opened, bare-silicon chip, no window, cooled “through the PCB” by a Peltier element; masking tape (emissivity 0.92) and the chip's thermal diode to calibrate the camera | DIAS PYROVIEW 380L, 8–14 µm | metal package: emissivity “~0.01”, mostly reflection |
| UC Riverside (Sadiqbatcha et al., 2019; Lu and Tan, 2024) [15] | laptop processors, an RTX 4060, a Coral TPU; cooled through the board | FLIR A325sc, 7.5–13 µm | through-PCB cooling “efficiency is significantly reduced” |
Lock-in thermography. The chip's power is “periodically amplitude-modulated”, which reveals heat sources “of a few µW corresponding to a local temperature modulation of a few µK”; it works for “backside inspection”, and at 1–25 Hz it sees “through 100-400 µm package material”. The same review notes that the mid-wave gives better spatial resolution [29]. This page's lock-in decay lengths (section 4) assume 5–10 Hz.
7.5 Esperanto's own public numbers
- Hot Chips 33 (2021). TSMC 7 nm, 570 mm²; “Power typically < 20 watts, can be adjusted for 10 to 60+ watts under SW control”; a 45 × 45 mm package with 2,494 balls and over 30,000 bumps; target cards “must be air-cooled”, and on a low-profile PCIe card the chip's budget “increases to ~60W” [3].
- IEEE Micro (2022). At “around 0.4 V, one chip would take about 20 W”; the Hot Chips performance numbers “were projections based on gate-level simulations” [4]. The 20 W is a modelled figure, not a measurement.
- The product page. The card has a “Full set of monitoring sensors” [5].
- Not found: any public junction-temperature limit, throttling or thermal-trip specification.
8. What our other pages already say
- The effect of overheating: at 600 MHz nothing limits the die on the cards' builds; the DRAM beside it has no thermometer and a fixed refresh; one swing to 120 °C costs 5.6–7.6 times the solder fatigue of a swing to 65 °C [R4]. A bare run would be such a swing, at best.
- Where the work sits: placement changes the time to the governor's 66 °C step. The same work took 1.62 times as long on aifoundry3's perimeter as in its interior (99% interval 1.51–1.73), and 2.04 times as long in short bursts on card 1, where the sustained test was insufficient; the host reads only the mean. A bare lid, spreading heat over mm, would not show where the work sits.
- The thermal-camera plan: removing a heatsink stays off shared cards without the lab admin (E-T20 in its full plan), since the bare lid is (nearly) isothermal (section 4); its view of the back of the board is the spatial channel that survives with the heatsink on.
- Spatial temperature: the chip's 35 sensors, and why the host sees one mean.
9. Method and caveats
Every number on this page comes from the cards' own record, the vendor's documents and firmware source, the owner's
thermal-camera plan, the models published with this page or the outside sources numbered at the end, and carries its kind where it
matters (the terms at the top). All of it is arithmetic on the record; no card was used. The outside sources were read, not reproduced, and their
caveats are listed with them. The model, nohs_calc.py, draws each uncertain
input from its range (uniformly, or log-uniformly for the resistances and the board's conductance) and reports the 5th,
50th and 95th percentiles; its seed is fixed, and its full printout is published beside it. Section 5 rests on two more
models built the same way: the envelope model, lowclock_calc.py (2,000 draws), which takes each card's
metered rails apart and re-runs the bare package's steady state, its window after the host boots and the heatsink's at
every clock and voltage policy, and the imaging model, imaging_calc.py (240 draws), a layered conduction
model of the die, lid, heatsink, substrate and board with the camera's and the sensors' noise. The measurement plan's
own predictions come from lc_plan_calc.py, which imports nohs_calc.py.
Which inputs are measured, and which are estimates
- Measured or fitted on our cards: each card's idle power law and idle point; card 0's power at 300 MHz
(12,035 samples at 60–64 °C,
card0_guard.py); aifoundry2's Foster chain; the heating times with the heatsink; the peak draw. - From the vendor: the package's dimensions, the card's 88 W input and its power tree.
- Estimates: θJB 1–3 °C/W and θJC 0.03–0.10 °C/W; the board's in-plane conductance; heat-transfer coefficients in still air and under a 1–2.5 m/s fan; the lid's emissivity; the share of idle power that is off the die (3.8–9 W by card) and the share of its slope on the die (0.80–0.95); the heat capacities of the package (9–18 J/K in the transient model) and of the nearby board (30–80 J/K); a room of 22–30 °C; and the 30–90 s host boot. The unmetered 13.7–17.5 W at 70 °C is split between die and board only by inference [R8]. The §3 transients draw the lid and board paths directly, an effective θ of about 3.8–7.8 °C/W in still air and 1.9–3.8 °C/W with a fan model.
- The laws are extrapolated. They were fitted up to 82–85 °C. Bare, the loop gain passes 1 inside that range, so the verdict does not rest on the extrapolation; the times to 120 °C, the powers at 116 °C, card 1's 75 W point near 99 °C, card 0's law above 64 °C and the with-heatsink threshold of 98.7 °C do.
- Section 5 adds: the three voltage policies, and what is safe at 400 MHz and for the NoC at 400 mV; one minion clock tree shared by the three cards (scaled by V²); leakage's fall with voltage, fitted to two anchors on different cards and firmware (aifoundry2's 800 MHz idle, card 0's 300 MHz rails), which pull against each other; the SRAM's, held only by its part of the 800 MHz step; the imaging model's die and lid thicknesses, paste resistances, emissivities, camera noise and its 4-fold drift allowance, an IR-window cooler's window and coolant transmission, and the stock sensor's dither; the blower's 3–6 m/s; the 20–60 s to set a low point after the host boots. The chip's behaviour below 300 MHz is not established at all.
What would change the answer. A bare θJA below θ85, the largest resistance at which an idle die settles at or below 85 °C: °C/W at 600 MHz, which only a heatsink provides; at 100 MHz with the floor voltages and the NoC lowered the bar rises to °C/W, within reach of a 3–6 m/s blower (1.8–3.0 °C/W) in steady state, though not through the plan's power-on gate, and not of still air or an ordinary fan model. A much flatter leakage law: card 0's at 300 MHz and 399 mV, where the card is marginal, or leakage falling with voltage more steeply than the envelope model has it, which Stage B of the measurement plan would measure model. A firmware limit that acts at 600 MHz, or an automatic power cut, which would make short bare runs survivable; with lock-in through a taped lid, a short run at a slow clock would also be informative (section 5.4).
Data and reproduction. Everything is in
docs/reports/data/2026-09-30-without-heatsink:
the model (nohs_calc.py, its printout nohs_calc.out), card 0's analysis
(card0_guard.py, card0_guard.out), section 5's models (lowclock_calc.py and
imaging_calc.py, each with its printout and JSON; lc_plan_calc.py and its printout), the firmware's
PLL tables and voltage limits decoded (pll_modes.py, pll_modes.out), the firmware notes
(LOWCLOCK-FIRMWARE.md), the measurement plan (LOWCLOCK-PLAN.md), the outside research behind
sources [44]–[53] with its quotes (LOWCLOCK-OUTSIDE.md), and this page's data
(make_page_data.py, page.json).
V=docs/reports/data/2026-09-30-without-heatsink
python3 $V/nohs_calc.py > $V/nohs_calc.out # the model, seed 20260929 (about 3 minutes)
python3 $V/card0_guard.py > $V/card0_guard.out # card 0's idle power at 300 MHz, from the guard samples
python3 $V/imaging_calc.py --json $V/imaging_calc.json > $V/imaging_calc.out # imaging model (about 8 minutes)
python3 $V/lowclock_calc.py --json $V/lowclock_calc.json > $V/lowclock_calc.out # envelope model, after imaging
python3 $V/lc_plan_calc.py > $V/lc_plan_calc.out # the plan's predictions, after the envelope model
python3 $V/pll_modes.py > $V/pll_modes.out # the firmware's PLL modes and voltage limits
python3 $V/make_page_data.py # page.json: checks the printouts, adds the trajectories
python3 scripts/build-report.py esperanto-without-heatsink $V/page.json docs/reports/2026-09-30-esperanto-without-heatsink.html
Sources
The ET-SoC-1's documents are in aifoundry-org/et-man; repository paths are in yaroslavvb/et-soc1-prototyping, with line numbers as of 30 September 2026.
- [R1] Esperanto, ET Preliminary Datasheet, Rev 1.0: Fig. 9-1 (p. 33), §8.1–8.2 and §9.1.
Used: the package's dimensions; the missing ratings. - [R2] Esperanto, ET-PCIe-Dev-Card-V3: pp. 2, 3 and 5, and its text on the input switch.
Used: the lid photo, the fan and auxiliary headers, the 88 W input, the hot-swap switch. - [R3] docs/findings/14-card-behaviour.md, lines 138, 148–151, 161–177, 204–229, 268, 333–340 and 385–388.
Used: the peak draw, the card table, the governor table and the PMIC alarm on each build, the protection at 600 MHz, the sensors the host sees, card 0 at 115–117 °C on 25 September. - [R4] docs/findings/05-claims.md, lines 497–499, 651–666 and 934–936.
Used: the power on the die, the effect-of-overheating conclusions, the idle laws at 116 °C. - [R5] docs/findings/11-thermal-model.md, lines 18 and 29–33.
Used: the Foster chain, its intercepts, the load in the heatsink trajectory. - [R6] docs/findings/12-heat-management.md, lines 20 and 28.
Used: 80 → 90 °C with the heatsink in 19–26 s (random normal) and 107–167 s (ones); the matmul's 27.2 W. - [R7] docs/findings/16-dvfs-and-leakage.md, lines 39 and 185–191.
Used: the metered rails at aifoundry2's idle point; its operating points, 600 MHz at 0.517 V, 700 at 0.568 and 800 at 0.618. - [R8] docs/findings/19-observability-and-the-unmetered.md, lines 36, 51–53 and 67.
Used: the unmetered power and its split, the regulator loss. - [R9] docs/reports/data/2026-09-28-overheating/analysis:
idle_vs_temp.txtlines 1–24 andrunaway.txtlines 1–4.
Used: each card's idle law and its fitted range; the loop gain with the heatsink. - [R10] docs/reports/data/2026-09-27-chip-diagram/research/facts-numbers.json, line 1356.
Used: the die's size and the shire spacing. - [R11] docs/getting-started.md, lines 287–288.
Used: power cycles are the lab admin's, on request. - [R12] What trips people up on the AI Foundry lab (25 and 27 September), requests SH1 and SH6, and problem H21.
Used: no console or out-of-band access on record; the requested on-site check of card 0's cooling. - [R13] Pointing a thermal camera at the ET-SoC-1 (the owner's plan), its instrument and experiments; and the full plan behind it,
plan.md, linked from the page's footer: §1, §3 target 9, E-T20, §7 pitfall 10 and §8 questions 4 and 6.
Used: the P3's band and rate; the lid decay lengths it estimated (target 9); E-T20 and its line on aifoundry1's cards; polyimide tape (pitfall 10); open questions 4 and 6; the heatsink as the state behind the 107 against 162–167 s scatter (§1). - [R14] docs/reports/data/2026-09-25-claims-v3/AMENDMENTS.md, lines 841–849.
Used: card 0 is excluded from the campaign. - [R15] docs/reports/data/2026-09-30-without-heatsink/LOWCLOCK-FIRMWARE.md: the firmware notes, each fact with its file and line in et-platform (BL2 0.18.0, 0.20.0 and 0.21.0, the PLL mode tables, the PMIC driver) or the manuals, with
pll_modes.pydecoding the tables (§0–§3, §6–§7).
Used: the minions' one step clock and its 33-shire mask; whatDM_CMD_SET_FREQUENCYaccepts, 100 MHz as mode 62, nothing below, and the minion-debug commands' reach; the dividers' headroom; each rail's accepted range, 0.18.0's missing check and flash write; the safe state and its TODO; the governor on each card; the regulator module. - [R16] docs/reports/data/2026-09-30-without-heatsink/LOWCLOCK-PLAN.md: the low-clock measurement plan (design only, not run), with its predictions from
lc_plan_calc.py; §0, §2.2, §3, §4.4, §5, §7–§9.
Used: the card and why; the stages, their card time and predictions; the safety rules and the bare step's gates; the blower's θJA and power-on odds; the Arrhenius forms' room for the clock tree (its section A); the event flood and the race between the clock command and the governor, why the command can interrupt the governor, and the race at the current source.
Outside sources
Numbered in order of first citation, each with what the page uses from it. Read in September 2026. Caveats: the Edmund Optics prices come from search-result listings, since Edmund's pages block automated fetching; the full Hamann et al. paper was not read, so IBM's camera band is not confirmed (its windows and a companion paper point to the mid-wave); the 20 W at 0.4 V figure for ET-SoC-1 is from a modelled study, not a measurement. Sources [44]–[53] were added with section 5 on 30 September, numbered in order of first citation there: [44] is a generic product brief, not this chip's PLL; [47] summarises TSMC's paper at second hand; [49]'s figures are read off a bar chart.
- [1] AMD, UltraScale and UltraScale+ FPGAs Packaging and Pinouts, UG575 v1.22 (10 March 2026), chapters 7, 10–11: docs.amd.com/v/u/en-US/ug575-ultrascale-pkg-pinout (PDF: docs.amd.com/api/khub/documents/GnuVZDIoTcrlIuu8Aok5XQ/content).
Used: Table 10-1’s θJA with no heatsink and its conditions; no thermocouple between package and heatsink; thermal paste, 20–50 psi, removal by twisting at 40–60 °C (Laird's advice for phase-change material, quoted in chapter 11) and the solvent wipe of chapter 11; chapter 7's warning that washing solvents can compromise the lid adhesive. - [2] Microchip, PolarFire FPGA Packaging and Pin Descriptions, UG0722, Table 9-1: ww1.microchip.com/downloads/aemDocuments/documents/FPGA/ProductDocuments/PackagingSpecifications/polarfire_fpga_packaging_and_pin_descriptions_ug0722_v12.pdf.
Used: the MPF500T-FCG1152’s θJA in still air, at 1.0 and at 2.5 m/s; the MPF300T in its 29 mm package, lidded (9.50 °C/W) and as a bare die (11.21 °C/W). - [3] D. Ditzel, Esperanto ET-SoC-1, Hot Chips 33 (2021), slides: hc33.hotchips.org/assets/program/conference/day2/HC2021.Esperanto.Dave_Ditzel.presentation.v1submitted.pdf.
Used: the process and die area, the power figures, the package, “must be air-cooled”, the ~60 W card budget; the minions' “OPERATING RANGE: 300 MHz TO 2 GHz” (slide 7). - [4] D. Ditzel et al., IEEE Micro (2022): www.esperanto.ai/wp-content/uploads/2022/05/Dave-IEEE-Micro.pdf.
Used: about 20 W at around 0.4 V, a modelled figure; the Hot Chips performance numbers as gate-level projections; the minions' sweet spot “between 300 and 500 mV”, also modelled. - [5] Esperanto, products page: www.esperanto.ai/products/.
Used: “Full set of monitoring sensors”. - [6] Tom’s Hardware (Pabst), “Hot Spot” (2001), pages 3–6: www.tomshardware.com/reviews/hot-spot,365-3.html (also 365-4, 365-5 and 365-6).
Used: the Athlon 1400, Athlon MP, Pentium 4 and Pentium III run with their heatsinks pulled; the Athlon MP's board protection, too slow at “only 1 degree/s”, and its death near 300 °C (page 5); both Athlons dead “within fractions of a second” (page 6). - [7] Intel, Atom Processor 330 Datasheet, 320528-003, §3.5 and §4.3: www.intel.la/content/dam/doc/datasheet/atom-330-datasheet.pdf.
Used: the trip at about 125 °C (§4.3, THRMTRIP#), leakage beyond its reach, and the supply off within 500 ms (§3.5). - [8] H. Amrouch and J. Henkel (KIT), “Lucid Infrared Thermography of Thermally-Constrained Processors”, ISLPED 2015: ces.itec.kit.edu/img/Lucid_Infrared_Thermography_of_Thermally_Constrained_Processors.pdf.
Used: the metal package’s emissivity of about 0.01; masking tape at 0.92; cooling through the PCB with a Peltier; the Alpha model above 200 °C bare. - [9] Intel, Thermal and Mechanical Design Guidelines, 318734-017, §3.4 and Appendix D: www.intel.pl/content/dam/doc/design-guide/core-2-e8000-e7000-pentium-e6000-e5000-celeron-e3000-guide.pdf.
Used: case temperature at the lid’s centre; the groove; the 36-gauge type T thermocouple soldered at 150 ± 3 °C. - [10] Infineon, application note on linear-mode operation and the safe operating diagram of MOSFETs (2017): www.infineon.com/dgdl/Infineon-ApplicationNote_Linear_Mode_Operation_Safe_Operation_Diagram_MOSFETs-AN-v01_00-EN.pdf?fileId=db3a30433e30e4bf013e3646e9381200.
Used: the definition of thermal instability. - [11] Taiwan Semiconductor, “Thermal Runaway on Schottky Diodes” (white paper, 2020): services.taiwansemi.com/storage/resources/white-paper-6/White-Paper_Thermal-Runaway-on-Schottky-Diodes_EN_20200901.pdf.
Used: the condition dPtot/dTj < 1/Rthj-a. - [12] A. Vassighi and M. Sachdev, IEEE Transactions on Device and Materials Reliability (2006): doi.org/10.1109/TDMR.2006.876577 (abstract via api.openalex.org/works/doi:10.1109/TDMR.2006.876577).
Used: leakage feedback and thermal runaway in chips under burn-in. - [13] G. Bhat, S. Gumussoy and U. Ogras, arXiv:2003.11081: arxiv.org/abs/2003.11081.
Used: an unstable system runs away; a stable point may exceed the operating limit. - [14] F. J. Mesa-Martínez et al., ISCA 2007: users.soe.ucsc.edu/~renau/docs/isca07.pdf.
Used: the Athlon 64’s 5.3% per 13 °C; the mineral-oil cooler, its camera and 100 W; silicon’s 55% transmittance; water rejected as not transparent. - [15] S. Sadiqbatcha et al., DATE 2019: past.date-conference.com/proceedings-archive/2019/pdf/0519.pdf; and Lu and Tan, MLCAD 2024: par.nsf.gov/servlets/purl/10542848.
Used: the FLIR A325sc at 7.5–13 µm, the chips imaged, the reduced efficiency of through-PCB cooling; that a thin spreader does not move the heat sources, which a 2D Laplacian recovers in simulation. - [16] DIAS Infrared, PYROVIEW 380L compact (the page now lists its successor, the 380L compact+): dias-infrared.com/products/infrared-cameras/infraredcameras-pyroview-compact.
Used: the 8–14 µm band, as listed for the current 380L compact+; KIT used the earlier 380L. - [17] Crystran, germanium: www.crystran.com/optical-materials/germanium-ge.
Used: the 8–14 µm band, the 53% reflection loss, opacity above about 350 K. - [18] Edmund Optics #26503, 50 mm × 1 mm germanium window, AR-coated for 8–12 µm: www.edmundoptics.com/p/50mm-dia-x-1mm-thick-8-12mum-ar-coated-ge-window/26503/.
Used: about $919, from a search-result listing (September 2026), not checked on the page. - [19] Specac, TN21-04, transmission windows: specac.com/wp-content/uploads/2022/06/TN21-04-Transmission-Windows.pdf.
Used: zinc selenide, barium and calcium fluoride, fused silica, KBr and NaCl. - [20] Edmund Optics #23151, 50 mm × 2 mm zinc-selenide window, coated for 8–12 µm: www.edmundoptics.com/p/50mm-dia-x-2mm-thickness-8-12mum-coated-znse-window/23151/.
Used: about $875, from a search-result listing (September 2026), not checked on the page. - [21] Crystran, zinc sulphide, FLIR grade: www.crystran.com/optical-materials/zinc-sulphide-zinc-sulfide-flir-zns/.
Used: 1.0–13 µm. - [22] Knight Optical, AMTIR: www.knightoptical.com/custom/infrared-optics/amtir/.
Used: AMTIR-1 at 0.75–14.0 µm. - [23] T. P. Mollart and K. L. Lewis, “The Infrared Optical Properties of CVD Diamond at Elevated Temperatures”, physica status solidi (a) 186 (2001): doi.org/10.1002/1521-396X(200108)186:2<309::AID-PSSA309>3.0.CO;2-I.
Used: CVD diamond for “some more extreme” 8–14 µm windows. - [24] EKSMA Optics, barium fluoride windows: eksmaoptics.com/optical-components/uv-and-ir-optics/barium-fluoride-baf2-windows/.
Used: the useful range 0.265–10 µm; $264 at 50.8 mm. - [25] Edmund Optics #7803, 50 mm uncoated calcium-fluoride window: www.edmundoptics.com/p/50mm-diameter-uncoated-calcium-fluoride-window/7803/.
Used: about $295, from a search-result listing (September 2026), not checked on the page. - [26] Crystran, sapphire: www.crystran.com/optical-materials/sapphire-al2o3.
Used: 0.17–5.5 µm. - [27] Crystran, silicon: www.crystran.com/optical-materials/silicon-si/.
Used: silicon as a 3–5 µm window; the oxygen band at 9 µm in Czochralski silicon. - [28] Topsil, HiTran application note (October 2013): www.topsil.com/wp-content/uploads/2023/05/hitran_application_note_october2013.pdf.
Used: the lattice absorption at 9 and 11–16 µm. - [29] O. Breitenstein et al., lock-in thermography, ASM 2011: www-old.mpi-halle.mpg.de/mpi/publi/pdf/10496_11.pdf.
Used: lock-in modulation, its µW and µK sensitivity, backside inspection through 100–400 µm at 1–25 Hz; silicon’s doping-dependent transmission; the mid-wave’s resolution; that heat sources the trigger does not affect do not appear in a lock-in image. - [30] Wikipedia, “Mulling (spectroscopy)”: en.wikipedia.org/wiki/Mulling_(spectroscopy).
Used: Fluorolube’s carbon–fluorine absorptions from 1300 to 400 cm⁻¹; a mull pressed between salt plates. - [31] Hom et al. (Stanford, with AMD), ITHERM 2012: nanoheat.stanford.edu/wp-content/uploads/2012/09/ITHERM2012_Lewis_FINAL.pdf.
Used: the Galden HT-170 plenum error and its opaque 2 mm film; FC-70’s mid-wave transmission, and that it too appeared opaque to their microscope; the sapphire window; the 3.4–5.1 µm band. - [32] Wikipedia, “Nujol”: en.wikipedia.org/wiki/Nujol.
Used: mineral oil’s major absorption peaks. - [33] International Crystal Laboratories, Fluorolube and Nujol: www.internationalcrystal.net/fluorolube-nujol/.
Used: Nujol from 1370 cm⁻¹ into the far infrared, with one absorption band at 720 cm⁻¹. - [34] Wikipedia, “Socket A”: en.wikipedia.org/wiki/Socket_A.
Used: Socket A dies were exposed, and their corners were susceptible to damage when coolers were installed incorrectly or systems handled roughly. - [35] nvidia-smi manual page: www.mankier.com/1/nvidia-smi.
Used: the GPU slowdown and shutdown temperatures. - [36] TechRadar, an AMD Ryzen 4000 APU running Crysis without a CPU cooler: www.techradar.com/news/so-an-amd-ryzen-4000-apu-can-apparently-run-crysis-without-a-cpu-cooler.
Used: the Ryzen 3 4300U run, which the article hedges as “apparently” achieved, its limit cut to 90 °C, its Cinebench score. - [37] Thermal Grizzly, on delidding Intel Arrow Lake processors: www.thermal-grizzly.com/en/blog/important-information-regarding-the-delidding-of-intel-arrow-lake-cpus.
Used: 165–180 °C for a soldered lid; the voided warranty. - [38] XDA Developers, “Why I will never delid my CPU”: www.xda-developers.com/why-i-will-never-delid-my-cpu/.
Used: one slip damages the surface-mounted parts. - [39] H. F. Hamann et al., IEEE Journal of Solid-State Circuits (2007): doi.org/10.1109/JSSC.2006.885064 (abstract via api.openalex.org/works/doi:10.1109/JSSC.2006.885064; the full paper was not read).
Used: the IR-transparent heat sink. Its camera band is not stated in the abstract; a companion paper reportedly used an InSb detector (from a search snippet, not read). - [40] US 7,167,806 B2 (IBM): patents.google.com/patent/US7167806B2/en.
Used: the coolants (including water or a cold gas), windows and duct; 200 W/cm² at a 70 °C rise. - [41] F. J. Mesa-Martínez, E. K. Ardestani and J. Renau, ASPLOS 2010: users.soe.ucsc.edu/~renau/docs/asplos10.pdf.
Used: the 3 mm sapphire window. - [42] K. Dev et al. (Reda group, Brown): arxiv.org/pdf/1808.09651.
Used: the SC5600 at 2.5–5.1 µm; the 1 mm oil channel between sapphire windows; oil as the thermal interface. - [43] US 10,175,705 B2 (Brown University): patents.google.com/patent/US10175705B2/en.
Used: oil at 5 m/s and 20 °C; about 90% of the heat upward. - [44] Movellus, HPDPLL product brief (TSMC 7 nm): anysilicon.com/wp-content/uploads/2020/06/Movellus-HPDPLL-Product-Brief3.pdf.
Used: “Output Frequency 4MHz - 6GHz”, for the vendor's generic PLL; the brief does not describe the ET-SoC-1's instance. - [45] R. G. Dreslinski, M. Wieckowski, D. Blaauw, D. Sylvester and T. Mudge, “Near-Threshold Computing: Reclaiming Moore's Law Through Energy Efficient Integrated Circuits”, Proceedings of the IEEE 98(2) (2010): doi.org/10.1109/JPROC.2009.2034764 (copy read: courses.grainger.illinois.edu/CS534/fa2021/reading_list/5a.pdf).
Used: SRAMs have an energy-optimal voltage “by approximately 100 mV” higher than processors'. - [46] Texas Instruments, TPSM831D31 product page: www.ti.com/product/TPSM831D31.
Used: “Output voltage range: 0.25 V to 1.52 V”, for the module that feeds the minion and NoC rails. - [47] Chipworks, “IEDM 2016 – Setting the Stage for 7/5 nm”, Solid State Technology (18 January 2017), a summary of S.-Y. Wu et al., IEDM 2016 (doi.org/10.1109/IEDM.2016.7838333, not read): sst.semiconductor-digest.com/chipworks_real_chips_blog/2017/01/18/iedm-2016-setting-the-stage-for-75-nm/.
Used: TSMC N7's “DIBL is ~40 mV/V”, a secondary source. - [48] E. Le Sueur and G. Heiser, “Dynamic Voltage and Frequency Scaling: The Laws of Diminishing Returns”, USENIX HotPower '10: www.trustworthy.systems/publications/nicta_full_text/4158.pdf.
Used: “P = Cf V² + Pstatic”: the clock cuts only the dynamic term. - [49] D. Ditzel, “Real World Results using Thousands of RISC-V Cores for AI and Beyond”, RISC-V Summit (13 December 2022), slides: hosted-files.sched.co/riscvsummit2022 (Esperanto Ditzel, Thousands of RISC-V Cores for AI and Beyond).
Used: slide 14's idle card, 20.0 W, with 5.8 W on the minions, 1.6 W on the NoC and 1.7 W on the SRAM, read off its bar chart; it gives no clock or temperature. - [50] G. Ruhl et al. (Intel), “An IA-32 Processor with a Wide Voltage Operating Range in 32nm CMOS”, Hot Chips 24 (2012): old.hotchips.org/wp-content/uploads/hc_archives/hc24/HC24-6-Tech-Scalability/HC24.29.625-IA-23-Wide-Ruhl-Intel_2012_NTV_iA.pdf.
Used: “1.1V/741MHz/445mW to 380mV/10MHz/1.5mW”; “Leakage power scales from 3% @1.1V to 50% @ 0.38V”. - [51] R. Cochran and S. Reda, “Spectral techniques for high-resolution thermal characterization with limited sensor data”, DAC 2009: doi.org/10.1145/1629911.1630037 (abstract via OpenAlex).
Used: a chip's thermal state reconstructed almost fully from a few sensors, on a 16-core processor. - [52] M. Kögel et al., “Lock-in Thermography for the Localization of Security Hard Blocks on SoC Devices”, ISTFA 2023, pp. 352–359: doi.org/10.31399/asm.cp.istfa2023p0352 (abstract: dl.asminternational.org/istfa/proceedings-abstract/ISTFA2023/84741/352/28617).
Used: firmware functions activated periodically, located “at the die level on a modern SoC”; the abstract does not name the SoC or say whether it was opened. - [53] S. Huth, O. Breitenstein, A. Huber, D. Dantz, U. Lambert and F. Altmann, “Lock-In IR-Thermography – A Novel Tool for Material and Device Characterization”, Solid State Phenomena 82–84 (2002): doi.org/10.4028/www.scientific.net/ssp.82-84.741 (copy: www-old.mpi-halle.mpg.de/mpi/publi/pdf/540_02.pdf).
Used: “35 µK (effective value) after 16 min”, with a cooled camera.
Related reports
- Pointing a thermal camera at the ET-SoC-1 — the owner's thermal-camera plan: what the camera can resolve, and where to point it.
- The effect of overheating — processor temperature limits, and what heat costs this chip.
- Where the work sits — placement and the time to the governor's 66 °C step (E52).
- Power and temperature telemetry — every way the card's power and temperature can be read, and the 34-shire voltage map.
- Heat per millimetre — what it costs to move a bit one millimetre across the chip's mesh.
- Spatial temperature — the 35 sensors, and why the host sees one mean.
- Limits of observability — the hub: every report, and what the meters cannot see.
- A compiler that knows where the fan is — a pitch (22 September), not a measurement report: the case for a thermally aware compiler, citing ET-SoC-1 measurements.
Versions. 30 September 2026: first edition, with a citation check of every outside source and a review of the page applied before publication. 30 September 2026, evening: the low-clock envelope added at the owner's request (section 5: how low the clock can go, what a slow card draws, whether it holds bare, what a camera would see, the alternatives ranked and a measurement plan), with the envelope and imaging models, the firmware notes and the plan published beside the page, outside sources [44]–[53], and the answers, the options and what would change the answer updated to match; sections 5–8 became 6–9.