{"data":{"items":[{"id":"00400882-4f84-4185-b996-e1bf30f963c2","excerpt":"\"Data center in a Box (on Wheels)\" 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks — I thought this would be relevant to the homelab subreddit so I'm adding it here, just to put the information out there and discuss if there is any interest. I am an IT infrastructure engineer","url":"https://www.reddit.com/r/homelab/comments/1veh5gt/data_center_in_a_box_on_wheels_256gb_vram512gb/","role":"demand","weight":1.1656945,"occurredAt":"2026-08-03T15:46:39.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"homelab","intent":"alternative_search","painScore":0.33095473,"sentiment":0.7525773,"confidence":0.87583333,"matchedPatterns":["alternative_to","manual_process"],"statement":"Not quite Nemotron or Deepseek level, but a very good \"lower cost\" alternative to its newer 5.0 versions.","title":"\"Data center in a Box (on Wheels)\" 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks","body":"I thought this would be relevant to the homelab subreddit so I'm adding it here, just to put the information out there and discuss if there is any interest. I am an IT infrastructure engineer by profession, so my contribution to the conversation is mainly from a hardware/systems perspective rather than from a Machine Learning researcher standpoint. I got my start with HPC's (Beowulf clusters) around ten years ago when I was a Physics undergrad in university, and this is what the experience has come to almost a decade later. Not everyone is going to want to read all of this, that's perfectly fine, the extras are just for those who want the info.\n\nStarting goal/idea:\n\nBuild an all-in-one creative design workstation to support a small business. This machine should be capable of effectively inferencing frontier MoE models; aiding the business in language/text tasks where English may not be everyone's native language. Additionally, it should be capable of simultaneous image generation tools for graphic design users, enabling rapid image editing and presentation tweaks for marketing, without the business ever having to worry about API credits or hard limits on tool usage. The idea is that a 3090 stack, which is still a generally \"good\" performer for LLMs, would be \"led\" by two 5090s to handle the heavy lifting of the visual creative work (one dedicated to image generation, one dedicated to image editing) to complement each other in a \"sweet spot\" on cost, raw performance, and creativity potential. This configuration also grants some flexibility to allocate a 5090 to the LLM stack for best prompt processing possible where desired. The end result would indicate that this goal has been achieved.\n\n# Overview\n\nSpecs\n\nCPU: 64 Core TR 3995WX\n\nRAM: 512Gb DDR4-3200 ECC\n\nVRAM: 256Gb GDDR6x/GDDR7 (8x3090's + 2x5090's)\n\nEnclosure: Core W200 Thermaltake Case\n\nMobo: ASUS Pro WRX80E-SAGE/SE Wifi\n\nPSU: 1300W+1600W (2900W combined), with OCP, linked via PSU2PSU\n\nStorage: 4Tb Nvme (fast) + 4Tb HDD (slow) + 8 or so 1Tb SATA SSDs (mid) over USB as needed\n\nOS: Ubuntu 25.10\n\nOther: 3 Bifurcation cards, 10 risers of various lengths\n\nFront end: Open WebUI\n\nBack end: llamacpp/koboldcpp\n\nIntended for (Recommend):\n\nLarge MoE inferencing, simultaneous LLM + ComfyUI (x2) operation, power users who may commonly hit credit limits, creative or technical professionals who can leverage these tools to compound productivity and complete objectives in shorter time.\n\nNot intended for (Do not recommend):\n\nTraining, multi-concurrent inferencing, performance maxing, extreme frontier model inferencing at high quants, casual users just looking for roleplay.\n\nResult summary:\n\nUsing the W200 as the platform for its generous real estate and configuration flexibility, all ten cards and components were able to find a permanent place in the enclosure without major concessions. The drive bay area was the only space that had to be completely repurposed for GPU mounting, and for us this was not a problem. The chamber with the cards hanging from the top is fairly hollow, so with the 140mm fan stack on the front and side there is a wind tunnel effect where the air blows in through the front and side, cooling the cards as it makes its way out the back/top. Depending on ambient temp, at idle the card with the highest temp usually hovers in mid to high 40s Celsius with the lowest in the mid 20's C (three 3090's are hybrids= fantastic for temperatures, but radiator mounting adds a logistical headache). When actively inferencing, the highest temp card may reach the mid 60s during sustained loads. Only when running image or video gen tasks will the 5090 running ComfyUI reach the 70's, but these are very brief intermittent workloads, so temperatures by our measurement has proved satisfactory over time. This result enables the small business to have full LLM, image generation (\\~9 seconds), and image editing (\\~8 seconds) capabilities on tap all on a single node so the data remains centralized, and provides much faster performance compared to the Cloud API they came from; in this case ChatGPT, where generation jobs could take 1+min, and has hard limitations. I just do not know how well this kind of setup would work with other vendor or card models; in a homogenous GPU cluster or one with notably less powerful image gen cards than the 5090, the performance would predictably be much lower.\n\nThings that surprised/stuck with me about the end result:\n\n* Noise. I expected this to sound like a jet taking off when operating, but that is not the case. It's a satisfying button click to come alive, then it's a low gentle hum going forward, nowhere near the kind of fan noises I'm used to hearing in server rooms. Even under load, the CPU 120mm radiator fans (exhausting out the top) are pretty much all I hear, the 140mm fans on front and sides I assume must be helping to contain the acoustics. I have built many gaming PCs over the years and own a top-tier gaming PC-- and I would not be able to distinguish this as any louder than those, especially at idle.\n* Utility. I planned for this to be used primarily for a small creative business, but what I did not expect was how I would find it so indispensable in my personal life as an IT professional. Being an infrastructure engineer, coding is not my wheelhouse. When I am the only IT staff on site or there is nobody else available to work with specific expertise like SQL, powershell/python scripting, or troubleshooting very specific/niche technologies, having this tool on standby I feel has paid itself over just within my career. It has helped me turn processes that may have otherwise took me hours into minutes, days into hours, even months into a matter of weeks/days. After using the tool extensively I hit a point where I had to acknowledge how local LLMs have moved definitively beyond being a toy or novelty; when deployed intelligently something like this can be a major asset for professional users.\n* Wheels. Sounds extremely minor, until you realize that no matter how happy the cards are with their individual temps: there are still ten high-power GPUs dumping heat into the room. That means unless you use a complex radiator solution or special venting to get heat outside, the room will get toasty and there is normally not a direct solution for this. The wheels however offer an indirect solution. Plan to work in the office that day? Wheel it into the guest bedroom and let it run over Wi-Fi. Plan to work away from home? Wheel it into the office, put it on LAN, and access it over a private VPN connection. If you can't stop the room from heating, then you can at least choose what room gets the heat, and as someone who has lived with computers extensively this is a hugely underrated perk.\n\nCaveats: To operate at its best, I recommend leaving the glass side panel off for improved airflow.\n\nTypical activity over a day:\n\nBoots up around 5:30am, start up the ComfyUI server(s), start loading a model, go get coffee, fully ready for use within 15-20 min. Shut down occurs usually around 8pm later in the day. Total daily activity, \\~12-14 hours.\n\n# Cost Breakdown\n\nLaying it out, because I know it will be asked, even though I am aware this is unfortunately not reproducible in the current market. Some components like the SSDs were acquired privately long before the RAM and hardware price hikes, so my timing getting certain things was extremely fortunate for the build budget. Some figures are exact, some are slightly rounded depending on if I found the original receipt.\n\n|Component|Qty|Source|Unit Cost|Subtotal|\n|:-|:-|:-|:-|:-|\n||||||\n|RTX 3090 24Gb|8|eBay|750-1000|6500|\n|RTX 5090 32Gb|2|Retail|2500-3000|5500|\n|TR 3995WX|1|eBay|1068.43|1068.43|\n|WRX80E-SAGE-SE|1|Amazon|949.99|949.99|\n|DDR4 ECC 64Gb|8|Amazon|81.99|695.28|\n|TT Core W200|1|Amazon|499.99|499.99|\n|PSU 1300/1600|2|Amazon|250-350|600|\n|4Tb nvme|1|Amazon|221.05|221.05|\n|1Tb SSD|8|Personal|60|600|\n|Risers (varying length)|10|Amazon|40-80|480|\n|Bifurcation cards|3|Amazon|50|150|\n|**Total**||||**\\~$17k**|\n\n# Problems/Stability Writeup\n\nThe Space Problem:\n\nProbably the first major hurdle in attempting something like this is figuring out, even theoretically, how to put 10 cards in a box in any kind of configuration that is not somehow detrimental to the hardware. I had considered modified mining rig frames at first, but I really wanted something with more robust rigidity in its structure, with breathability, and allows some degree of portability. There are unfortunately not a lot of options for configurations like what I was imagining; I had looked into various cabinets and extended tower cases, but the dual full tower chamber design of the W200 was the only one where I could see this idea potentially working. I'm certain other solutions probably exist, maybe even some that allow mobility, but the W200 was really the best option I could find that checked the boxes of enclosure, space real estate, high air throughput, and semi portability. I recommend the W200 to solve the space problem, assuming it is available to you.\n\nThe Bifurcation Problem:\n\nAmong the other hurdles you may run into in assembling something like this may involve bifurcation cards. The cards rely on specific BIOS settings for things to work correctly, and if these settings are not put in place **before** everything is connected you may either see no output like the system is hanging or cards just won't show up once in the OS. Start with one GPU in a slot, no bifurcators yet; go into BIOS, and manually set each slot that will be split to bifurcation mode. While here, ensure above 4G decoding is enabled, Resizable BAR enabled, and SR-IOV enabled, this has given me best stable configuration with Ubuntu and multiple GPUs. If you use risers, especially if they are mixed generations, I highly recommend setting the Gen and lane speeds for each PCIe slot in the BIOS manually to ensure the system can effectively communicate with each card. Optimize riser Gen/speeds to be roughly similar to keep one card from dropping to a slower rate than the others--this does not necessarily impact inference performance as much as it heavily impacts model load time. No, you may not have any card running at the fastest possible Gen bandwidth at all times with this config, but loading a 200+gb model over an averaged Gen 3/4 x8/x16 PCIe speed will often be noticeably faster than if you let the system decide to make one or multiple cards run at Gen 1 x1.\n\nThe Power \"Problem\":\n\nPower and heat concerns I think remain to be among the biggest sources of skepticism regarding this project so I think it deserves a section here. To be fair, the concern in most situations would be understandable. If all ten of these cards pulled at or near their full TDP for sustained periods, components would melt. Fires would start. Neighbors would be asking awkward questions. However in reality, only 1400-1600W of the 2900W PSU capacity gets utilized under sustained load, and inter-GPU bandwidth bottlenecks are what allows this. In a way it is like a natural regulator that ensures the cards remain power restrained, and it is just physics, no voodoo necessary. When MoE's are sharded across a GPU stack, each forward pass requires all communication over PCIe, so the GPUs spend more time waiting on information from the last GPU than actually crunching compute. This means instead of needing to handle thousands of Watts to feed all the components running at full blast, it is a much more manageable 1400-1600W under LLM operation which can comfortably fit on a 20A/120V circuit (2400W max). On a per-GPU basis this may sound inefficient since the individual cards are being \"underpowered\", but this could arguably be flipped as being highly efficient on a per-node basis (\\~1600W sustained versus 4500W+ if all cards were \"fully\" utilized). As a precaution, I may set a power limit on the 3090's to 200W and the lead 5090 to 400W, but in practice the 3090's only pul","offTopic":true},{"id":"1c4f40d2-63d1-498a-82dd-402844e787e5","excerpt":"Gaming server finally complete :) — Hello Community!\n\nI hope I chosen correct sub-reddit as this is combination of Gaming, Bazzite, Proxmox, PC Builds and crazy DIYing. So let me present to you my Gaming Server. Here are the specs:\n\n    root@void:~# inxi -Fxz  \n    System:\n      Kernel: 6.11.11-2-pve arch: x86_64 bits:","url":"https://www.reddit.com/r/homelab/comments/1qmoe3v/gaming_server_finally_complete/","role":"request","weight":1.0865216,"occurredAt":"2026-01-25T16:53:02.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"homelab","intent":"feature_request","painScore":0.28455752,"sentiment":1,"confidence":0.84583336,"matchedPatterns":["wish","missing_feature"],"statement":"To pull it off I had to configure PRIME Z790-P WIFI D4 to bifurcate v5.0 x16 slot to two x8 x16 slots, many cases lack space to do so, until I found one ASUS PRIME AP303.","title":"Gaming server finally complete :)","body":"Hello Community!\n\nI hope I chosen correct sub-reddit as this is combination of Gaming, Bazzite, Proxmox, PC Builds and crazy DIYing. So let me present to you my Gaming Server. Here are the specs:\n\n    root@void:~# inxi -Fxz  \n    System:\n      Kernel: 6.11.11-2-pve arch: x86_64 bits: 64 compiler: gcc v: 12.2.0\n      Console: pty pts/0 Distro: Debian GNU/Linux 13 (trixie)\n    Machine:\n      Type: Desktop System: ASUS product: N/A v: N/A serial: N/A\n      Mobo: ASUSTeK model: PRIME Z790-P WIFI D4 v: Rev 1.xx serial: <filter>\n        UEFI: American Megatrends v: 0809 date: 01/06/2023\n    CPU:\n      Info: 24-core (8-mt/16-st) model: 13th Gen Intel Core i9-13900KF bits: 64 type: MST AMCP\n        arch: Raptor Lake rev: 1 cache: L1: 2.1 MiB L2: 32 MiB L3: 36 MiB\n      Speed (MHz): avg: 800 min/max: 800/5500:5800:4300 cores: 1: 800 2: 800 3: 800 4: 800 5: 800\n        6: 800 7: 800 8: 800 9: 800 10: 800 11: 800 12: 800 13: 800 14: 800 15: 800 16: 800 17: 800\n        18: 800 19: 800 20: 800 21: 800 22: 800 23: 800 24: 800 25: 800 26: 800 27: 800 28: 800\n        29: 800 30: 800 31: 800 32: 800 bogomips: 191692\n      Flags: avx avx2 ht lm nx pae sse sse2 sse3 sse4_1 sse4_2 ssse3 vmx\n    Graphics:\n      Device-1: Advanced Micro Devices [AMD/ATI] Navi 22 [Radeon RX 6700/6700 XT/6750 XT /\n        6800M/6850M XT] vendor: Sapphire driver: vfio-pci v: N/A arch: RDNA-2 bus-ID: 0000:03:00.0\n      Device-2: Advanced Micro Devices [AMD/ATI] Navi 22 [Radeon RX 6700/6700 XT/6750 XT /\n        6800M/6850M XT] vendor: Sapphire driver: vfio-pci v: N/A arch: RDNA-2 bus-ID: 0000:06:00.0\n      Display: unspecified server: N/A driver: N/A tty: 195x54\n      API: EGL v: 1.5 drivers: swrast platforms: active: surfaceless,device\n        inactive: gbm,wayland,x11\n      API: OpenGL v: 4.5 vendor: mesa v: 25.0.7-2 note: console (EGL sourced) renderer: llvmpipe\n        (LLVM 19.1.7 256 bits)\n      Info: Tools: api: eglinfo,glxinfo x11: xdriinfo, xdpyinfo, xprop, xrandr\n    Audio:\n      Device-1: Intel Raptor Lake High Definition Audio vendor: ASUSTeK driver: N/A\n        bus-ID: 0000:00:1f.3\n      Device-2: Advanced Micro Devices [AMD/ATI] Navi 21/23 HDMI/DP Audio driver: vfio-pci\n        bus-ID: 0000:03:00.1\n      Device-3: Advanced Micro Devices [AMD/ATI] Navi 21/23 HDMI/DP Audio driver: vfio-pci\n        bus-ID: 0000:06:00.1\n      API: ALSA v: k6.11.11-2-pve status: kernel-api\n    Network:\n      Device-1: Intel Raptor Lake-S PCH CNVi WiFi driver: iwlwifi v: kernel bus-ID: 0000:00:14.3\n      IF: wlp0s20f3 state: down mac: <filter>\n      Device-2: Realtek RTL8125 2.5GbE vendor: ASUSTeK driver: r8169 v: kernel port: 4000\n        bus-ID: 0000:09:00.0\n      IF: eno1 state: up speed: 1000 Mbps duplex: full mac: <filter>\n    Bluetooth:\n      Device-1: Intel AX201 Bluetooth driver: N/A type: USB bus-ID: 1-14:5\n      Report: This feature requires one of these tools: hciconfig/bt-adapter\n    RAID:\n      Hardware-1: Intel Volume Management Device NVMe RAID Controller Intel driver: vmd v: 0.6\n        bus-ID: 0000:00:0e.0\n      Device-1: primary type: zfs status: ONLINE level: linear raw: size: 2.79 TiB free: 2.54 TiB\n        zfs-fs: size: 2.7 TiB free: 2.45 TiB\n      Components: Online: 1: nvme0n1 2: nvme1n1 3: nvme2n1\n    Drives:\n      Local Storage: total: 8.36 TiB lvm-free: 13.88 GiB used: 3.61 TiB (43.2%)\n      ID-1: /dev/nvme0n1 vendor: Kingston model: SKC3000S1024G size: 953.87 GiB temp: 29.9 C\n      ID-2: /dev/nvme1n1 vendor: Kingston model: SKC3000S1024G size: 953.87 GiB temp: 32.9 C\n      ID-3: /dev/nvme2n1 vendor: Kingston model: SKC3000S1024G size: 953.87 GiB temp: 27.9 C\n      ID-4: /dev/sda vendor: Patriot model: Burst Elite 120GB size: 111.79 GiB\n      ID-5: /dev/sdb vendor: Samsung model: SSD 870 EVO 2TB size: 1.82 TiB\n      ID-6: /dev/sdc vendor: Seagate model: ST4000DM004-2CV104 size: 3.64 TiB\n    Partition:\n      ID-1: / size: 36.93 GiB used: 23.49 GiB (63.6%) fs: ext4 dev: /dev/dm-1 mapped: pve-root\n      ID-2: /boot/efi size: 511 MiB used: 8.8 MiB (1.7%) fs: vfat dev: /dev/sda2\n    Swap:\n      ID-1: swap-1 type: partition size: 8 GiB used: 0 KiB (0.0%) dev: /dev/dm-0 mapped: pve-swap\n    Sensors:\n      System Temperatures: cpu: 32.5 C mobo: 33.0 C\n      Fan Speeds (rpm): fan-1: 648 fan-2: 509 fan-3: 0 fan-4: 0 fan-5: 0 fan-6: 0\n    Info:\n      Memory: total: 96 GiB available: 94.08 GiB used: 4.45 GiB (4.7%)\n      Processes: 696 Uptime: 3h 14m Init: systemd\n      Packages: 1287 Compilers: gcc: 14.2.0 Shell: Bash v: 5.2.37 inxi: 3.3.38\n\nOn the server there are two Gaming VMs powered by Bazzite using GPU passtrough of AMD 6750 XT from Sapphire, running with v4.0 x8 speeds. To pull it off I had to configure PRIME Z790-P WIFI D4 to bifurcate v5.0 x16 slot to two x8 x16 slots, many cases lack space to do so, until I found one ASUS PRIME AP303. This case allows placement of PSU in the front allowing you to gain space below motherboard for fans, in this case for additional GPU.\n\nParts used:\n\n* DIY Mess:\n   * PCIE X16 to X16 Riser Card Adapter PCI Express 3.0 16X 90 Degree Reverse Male to Female Converter Expansion Card\n   * PCIe4.0 Riser cable GEN4 for GPU\n   * PCI-Express 4.0 3.0 x16 1 to 2 Expansion Card Gen4 Split Card PCIe-Bifurcation x16 to x8x8 20mm Spaced Slots\n   * Threaded rods for folding GPUs and nuts\n* GPUs: 2x AMD 6750 XT\n* Case: ASUS PRIME AP303\n\nThen I simply drilled bigger holes for GPU mount brackets (only larger size was available in local HW store), tighten nuts on both sides, attached GPUs - disconnected from Motherboard, enabled bifurcation in bios, connected x16 splitter to the motherboard and prayed.\n\n  \nBut why?\n\nIn Home Assistant there's proxmox integration that allows me and my gf to power up those VMs, then we connect to them via Moonlight + Sunshine from our laptops and sometimes I hook it up to Nokia Streaming Box 8010 + cheap Chinese Projector with PS5 controller so that I can play from couch allowing both of us to play our favorite games. \n\nCase review:\n\n* Pros:\n   * Great Airflow\n   * Cool simple and minimal design\n   * Loot of options for fan placement\n   * Place for 1x3.5\" HDD\n* Cons\n   * Clipings for panels are plastic, they break, that's why they include it in the box with Case\n   * Included wiring for HDD/CPU load is hardwired my motherboard would need classic split connectors in order to light up power button\n\nWhat's next:\n\n* Once I get a 3D printer I would like to build a better stand for GPUs so that lower GPU is better supported and there's wider airflow\n* Use simple short riser cable instead of this long one which I had to twist and turn\n\nWhat I wish I've done better:\n\n* Use 3D print for hodling GPUs and have stronger PCIe slot mounting\n* Used AMD processor instead of Intel one, mainly cause there's easier upgrade path due to AMD4, AMD5 sockets\n* Bought smaller Motherboard with PCIe v5.0 and Bifurcation option\n* Before buying motherboard - checked orientation of SATA ports and what lines and with which speeds are supported#\n* Bought CPU with GPU - so that I can get to Bios even without a GPU - maybe possible even now but I think I disabled onboard graphics for some reason :shrug:.\n* This setup only allows v4.0 GPUs to be used in 32GB/s - newer GPUs have higher troughput - not sure how that would affect speeds.\n\nI wish there was a motherboard, that would support at least x8 on two x16 format slots and be in Mini-ITX format with slots like 1x x8 v4.0, 1x x16 v5.0 that would allow me to plug in 10G Ethernet adapter and two GPUs (maybe). Also I used this setup instead of Oculink as this solution was bit cheaper in the end - I will consider it in future builds.\n\nOverall I'm really happy with the setup as now I can utilize those GPUs better even when I had a GPU in x4 speed it was working fine.\n\nEdit 1: Added Why section","offTopic":true},{"id":"38b9cfc2-549e-41ee-8b4c-2cebd315ad83","excerpt":"Mac | Cubix | V620 | Ubuntu | ROCm | vLLM | Local AI Data Center — What a loaded title.\n\nIt started with the 2019 Mac Pro, however it has since grown into so much more, evolving from niche to explicitly unique. Allow me to explain.\n\n# TL;DR\n\n* Three 2019 Mac Pro systems (MacPro7,1)\n* Cubix Xpander Rackmount (8 PCIe slo","url":"https://www.reddit.com/r/MacPro2019LocalAI/comments/1v2u5k6/mac_cubix_v620_ubuntu_rocm_vllm_local_ai_data/","role":"request","weight":0.99997145,"occurredAt":"2026-07-21T20:24:27.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"MacPro2019LocalAI","intent":"feature_request","painScore":0.36,"sentiment":0.14942528,"confidence":0.7352731,"matchedPatterns":["missing_feature"],"statement":"I appreciate Rhino Technology, perhaps not for the missing rhino prop, but certainly for their communication and respect.","title":"Mac | Cubix | V620 | Ubuntu | ROCm | vLLM | Local AI Data Center","body":"What a loaded title.\n\nIt started with the 2019 Mac Pro, however it has since grown into so much more, evolving from niche to explicitly unique. Allow me to explain.\n\n# TL;DR\n\n* Three 2019 Mac Pro systems (MacPro7,1)\n* Cubix Xpander Rackmount (8 PCIe slots, passive cooling)\n* AMD Radeon PRO V620\n* AMD Radeon PRO W6900X\n* AMD Radeon PRO W6800X Duo\n* AMD Radeon PRO W6800\n* Sonnet eGPU Breakaway Box 750/750ex\n* Ubuntu Server 24.04 LTS — bare metal\n* ROCm 7.2.3\n* vLLM ~~0.24.0~~ 0.25.1\n* FP16 and AWQ\n* Qwen3.6-27B / gemma-4-31B-it\n* Several Hermes agents\n* \\[ SUCCESS \\]\n*  بانتظار أسمع منكم جميعًا\n\n---\n\n# The Dream\n\nAchieving the dream is the goal here. The journey is half the dream, with the technical goal being the ability to run 30 to 50 concurrent agents. Currently, that means Hermes agents, each with a unique profile, role, and human name. (Adam, Samar, Sami, Dalia, Basil, Leen, Ziyad, Sultan, and many more)\n\nOn this journey, I hope to master vLLM, multi-GPU setups, high concurrency, general optimization, and troubleshooting wherever possible.\n\nI will keep the actual goal and final purpose of all of this private for now.\n\n---\n\n# GPUs | AMD? | NVIDIA? | Tenstorrent?\n\nI previously discussed [multiple-GPU setups in this post](https://www.reddit.com/r/MacPro2019LocalAI/comments/1tpjj2m/mac_pro_2019_160_gb_vram_achieved_five_amd_gpus/). u/Guanaalex introduced me to the world of Cubix Xpanders, and I was hooked. I managed to find a 4U Cubix Xpander Rackmount on eBay. The seller was kind enough to offer it at a price I could reasonably afford. [Please support the seller, Mara7Electronics](https://www.ebay.com/str/mara7electronics).\n\nI decided to buy a full-fledged 42U server rack to host it and migrate all my hardware into it.\n\nI had previously bought a nice, rack-mountable [online double-conversion 3.6 kVA / 3.6 kW UPS](https://www.tecnoware.com/en-US/Dettaglio/Articolo/74) to power [the two Macs I was using](https://www.reddit.com/r/macpro/comments/1q9xeov/guide_mac_pro_2019_macpro71_w_linux_local_llmai/). I decided to buy a couple more: one for each Mac and one for the Cubix Xpander. I also decided to replace my daily-driver 2019 Mac Pro with a Mac mini M4, allowing the Mac Pro to become my third Local AI system: LinuxAI-03.\n\nAlthough I already had three AMD Radeon PRO W6800 GPUs that I had purchased for use as eGPUs, that plan was abandoned in favor of the Cubix Xpander's cleaner eight-GPU setup.\n\nI considered purchasing five more W6800 GPUs, eight AMD Radeon PRO AI R9700 GPUs, or even eight Tenstorrent Blackhole p150 AI accelerators. I considered NVIDIA GPUs for a quick second, but the cost quickly killed that idea. Eventually, I stumbled across AMD Radeon PRO V620 cards on eBay, which came with fan shrouds, had been flashed with W6800 firmware, and included a comment explaining that the V620 firmware could be restored for pure compute use.\n\nI had not considered these cards before. I barely knew anything about them. I looked them up and found them on eBay for a pleasant $350 USD each. Eight of them would cost about the same as three W6800 GPUs. The only challenge was cooling.\n\nLo and behold, the Cubix Xpander I had bought happened to be the model that supports passively cooled hardware. I did not give it another thought. I immediately started [discussions with the seller](https://www.ebay.com/str/rhinotechnologygroup). They refused to gift me a rhino prop with my purchase. I was kind of disappointed. I appreciate [Rhino Technology](https://www.ebay.com/str/rhinotechnologygroup), perhaps not for the missing rhino prop, but certainly for their communication and respect. Please support them.\n\nYou may notice that I did not consider Intel cards. The reason was simple: I did not know Intel's direction for its GPU business, and I did not want to invest in the hardware only to see development of its software stack discontinued if Intel sold or shut down that part of the business.\n\n---\n\n# The Data Center\n\nAlthough I had an old 12U server rack, it was more of a wall-mounted networking rack, and it was already full. I searched online for the 42U server rack I wanted, but everything was either moderately priced with no description beyond “42U,” or fully documented but insanely expensive.\n\nI ended up sending my son to the local computer market, which is labeled a bazaar even though it is not really one. I loved the experience for him. He managed to find several shops carrying server racks with the specifications I wanted. He then found the cheapest shop that also offered delivery and installation, and bargained with the shop owner.\n\nWith that, I had my first 42U server rack: front-to-back airflow, double mesh doors on both sides, and fans preinstalled at the top. The server rack was delivered and installed on the same day.\n\nNext came the UPS devices.\n\nThe Tecnoware UPS I mentioned earlier was no longer available for sale anywhere. Nothing online was both good enough and cheap enough. I sent my son back to the computer bazaar, but he could not find anything reasonably comparable to the UPS I already had in terms of its kilowatt-to-price ratio, online double-conversion capability, and rack-mountable design.\n\nI ended up searching [Haraj](https://haraj.com.sa), the local equivalent of Craigslist, for UPS options, as well as Microless, which I would describe as Dubai's version of Newegg. I found a local vendor selling enterprise-grade 6 kVA / 6 kW UPS devices from a well-known international manufacturer for roughly half price. The catch? They were old stock from mid-2023, apparently unsold hardware left over from a project whose contract had ended.\n\nI tried to purchase only two UPS units, but the company insisted on selling each one with three rack mounted battery packs and would not budge on the price. I was about to cancel the purchase when work pulled me away. Later, I had a nice conversation with u/Long-Shine-3701, who convinced me to go for it, particularly with my future green-energy project in mind.\n\nAt the time, I did not know exactly how old the batteries were. I only knew they were “old” and had generally been kept in room-temperature storage. Regardless, my goal was never to keep the servers alive for long periods during power outages. My main goals were to provide clean, pure sine-wave power and allow for safe shutdowns. It is worth noting that each battery pack contains twenty standard, replaceable 9 Ah battery cells, although I do not have the faintest idea how to replace them yet.\n\nI reached an agreement with the company to provide each UPS with four batteries, the maximum number supported by these UPS units, along with a warranty, free delivery, and installation.\n\nI went for it.\n\nI did the rack-space math. It went something like this: a 1U UPS plus four 3U battery packs, with 1U of space between each unit to reduce heat buildup and prolong battery life... Thirty-seven rack units?! That was almost my entire rack.\n\nI measured the data room quickly, then proceeded with a quick phone call to the server-rack supplier my son had found, followed by a bank transfer, and I had same-day delivery and installation of a second rack. I barely had 2 cm, roughly half an inch, of clearance after installing the second rack. It was a perfect fit. I felt like a child at a candy store at that point.\n\nThe next day, the UPS units and battery packs were installed. The company was concerned about the available power, but I had already purchased five 10 mm² copper conductors, obtained a second meter from the electric company for this setup, and purchased a couple of breakers—one manual and one smart—as well as power-distribution equipment.\n\nAll that remained was to hire an electrician to connect the second meter to the breaker in the room. I had already arranged for one to work on a Saturday so the task could be completed quickly. The plan was ready; only the execution remained.\n\nThe company set up the UPS units and battery packs and initially connected them to my home meter to charge the batteries and test the system. Everything seemed to be working well, pending grounding, neutral wiring, and connection of the second meter. If the absence of neural wiring questions for you, I used two live wires to complete the circuit, and obtain the higher voltage; 220 V rather than 110 V.\n\nThe electrical work, while impressive in my opinion, does not get a detailed mention here beyond the fact that it is now part of the home data center and is controlled through Home Assistant, after the electrician completed the connection. If anyone wants to know more, I would be more than happy to share.\n\n---\n\n# Resources\n\nWhile working on this project, I experimented and learned a great deal. I then shared a great deal and received a tremendous amount of valuable knowledge and education from the community, which changed my plans midway through the project.\n\nThe target was always higher concurrency through more VRAM. Unified memory, or uRAM, was not an option for me, as one of my goals was to master dedicated hardware—AI accelerators in one form or another—for inference.\n\nThe first idea was to add four eGPUs to the Mac that already had four GPUs. I bought:\n\n* ~~Four~~ Three AMD Radeon PRO W6800 GPUs. The fourth was canceled by the seller.\n* Four Sonnet eGPU Breakaway Box 750/750ex enclosures.\n\nThen the plan shifted to the Cubix Xpander, and I bought:\n\n* The Cubix Xpander\n* Eight AMD Radeon PRO V620 GPUs\n* A Mac mini M4 to replace my daily-driver 2019 Mac Pro\n* A fifth Sonnet eGPU Breakaway Box 750ex to use a PCIe card from the Mac Pro with the Mac mini\n* Two 42U server racks\n* Two enterprise-grade UPS units with four battery packs each\n* Two patch panels, one for each server rack\n* Two SilverStone HELA 2050R Platinum PSUs\n\nI then found a pair of Cubix Xpander Desktop Elite systems, each with four PCIe slots, and bought those as well.\n\nWith international shipping and double taxation, I have severely exceeded my budget. I have had to bring all further spending to a complete stop and limit myself to covering only operational and maintenance costs.\n\nThe electricity bill alone will be an insane operation expense.\n\nSomething worth mentioning though, I would love to get my hands on sixteen Tenstorrent Blackhole p150a-series accelerators and QSFP-DD 800G cables. Testing all of them on a single server using every available Cubix Xpander would truly push every piece of hardware involved to its limit. Had I possessed the necessary capital, that is probably the direction I would have taken instead. I am just putting the thought out there. A Tenstorrent Galaxy Blackhole or four would be insane as well, would it not? A guy can only dream.\n\nI am genuinely hopeful, believing in the work these guys are doing there. I would also like to highlight [Tenstorrent's documentation](https://docs.tenstorrent.com/) and [software stack](https://tenstorrent.com/developers).\n\n---\n\n# The Challenge\n\nI am happy to say that I am satisfied with the results, and I look forward to continue pushing further and expanding the stack.\n\n**Power:**\n\nThe first hurdle was power. Not its availability, but its deliverability.\n\nThe PSUs in the Cubix Xpander were only designed to power eight cards using 8-pin and 6-pin connectors. For the V620 cards, I had to replace those PSUs with SilverStone HELA 2050R Platinum units to provide dual 8-pin connections to each GPU. That is sixteen 8-pin connections total, at 150 watts each.\n\nThey cost me a pretty penny, but I was lucky enough to find them on [Microless](https://saudi.microless.com/) for half the price listed on Amazon and eBay.\n\n**Assembly:**\n\nDuring my first exploratory disassembly of the Cubix Xpander, I may have overtightened the screws. When it came time to open the unit again, install the new PSUs, and then install the GPUs, the screws simply would not budge. I was unable to open it. I even stripped the screw heads while trying...\n\nI performed some clever analysis and concluded that when I first opened the Cubix Xpand","offTopic":true},{"id":"72902d4b-d9a3-4ed8-9344-3d3fcbd3320c","excerpt":"My Very Distracting 15,000+ GIF PC — When I set out to create this computer, I wanted to build something I would truly be proud of — something that, at least from my perspective, felt unique and would hopefully last me as long as its predecessor did(9 years).\n\nI had initially “finished” this build using a large number ","url":"https://www.reddit.com/r/pcmasterrace/comments/1teuodq/my_very_distracting_15000_gif_pc/","role":"request","weight":0.83386666,"occurredAt":"2026-05-16T14:07:30.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"pcmasterrace","intent":"problem_report","painScore":0.18,"sentiment":0.27272728,"confidence":0.70666665,"matchedPatterns":["i_need"],"statement":"Whenever I need a quick break from gaming, or just need some brief entertainment, I can just look at my pc.","title":"My Very Distracting 15,000+ GIF PC","body":"When I set out to create this computer, I wanted to build something I would truly be proud of — something that, at least from my perspective, felt unique and would hopefully last me as long as its predecessor did(9 years).\n\nI had initially “finished” this build using a large number of Lian Li screen fans with GIFs on every fan screen, but soon after I completed that, I envisioned something much more ambitious, which is what I’m sharing here today.  I have still procrastinated on selling those expensive fans :/\n\nOver many weekends and evenings, I spent countless hours scouring the internet and downloading well over 17,000 GIFs throughout both the original fan-screen phase of the project and the much larger display phase that followed.  Even after scrapping a large number of GIFs during editing and organization, the final build still contains over 15,000 GIFs.\n\nIt takes 13 and a half hours to watch every display tile from beginning to end.\n\nI would download  several thousand gifs at a time from various sites, and then I used a random naming software to randomize all of the file names of what I downloaded, then then I'd use the same software to sequentially rename them using numbers. That's how I kept track of everything, and made sure gifs from the same topic didnt get clumped together.\n\nOn top of collecting all of those GIFs, I also spent roughly 200 hours video editing alone.\nI edited every single GIF to better fit the aspect ratio of its respective display tile, adjusted playback speeds to make clips appear as close to real-time as possible when needed, looped gifs that needed looping, and trimmed many clips down to focus only on their most important part - I wanted my gifs to be very sporadic.\n\nNearly every screen tile runs for 5 minutes before looping.  The only exceptions are the two 5-inch displays that have two tiles, which run for approximately 10 minutes and 15 minutes respectively.  Those two screens ended up configured that way because two additional displays that were originally intended to sit in front of the GPU stopped working before final installation; they had narrower aspect ratios for their tiles.  At that point, I decided to compromise rather than  buy replacements or sink even more time into additional editing work.  \n\nThe four larger screens are all run by three Raspberry Pis, all within the case.  All the Raspberry Pis and screens power on and off with the pc.  The Raspberry Pi 5s are supposed to be able to run two 4k displays at once, but I had issues running a second display with the 3840x2400 14.5” display, so I had to get an additional Raspberry Pi.  The remaining 9 displays all run their videos on their displays from a microsd card. No display cables needed, only usb power required.  No desktop processing power used towards these screens. Also, the bar display to the left of the larger vertical bar display is on a hinge and is held in place by a magnet.  I made it this way so that I can easily see any codes from the motherboard if needed in case any issues arise later on.\n\nWhenever I need a quick break from gaming, or just need some brief entertainment, I can just look at my pc.\n\nPC specs:\n\n• CPU — AMD Ryzen 9 9950X3D\n\n• Motherboard — MSI X870E Carbon WiFi\n\n• GPU — MSI Ventus 3X GeForce RTX 5090 OC Edition\n\n• RAM — 96GB DDR5\n\n• Storage — Samsung 9100 Pro 4TB NVMe SSD\n\n• Case — Lian Li O11 Dynamic Evo XL \n\n4 120mm noctua fans on the top\n2 120mm noctua fans on the side mounted on the exterior, no space inside.\n3 140mm noctua fans on the bottom\n6 140mm noctua fans on radiator push/pull\n\nDisplay setup:\n\n• 5× WOWNOVA 5\" USB displays (800×480)\n\n• 4× WOWNOVA 8.8\" USB displays (1920×480)\n\n• 1× MAGICRAVEN 14.5\" 4K portable monitor (3840×2400)\n\n• 1× VSDISPLAY 12.7\" stretched bar display (2880×864)\n\n• 2× Wisecoco 14\" 4K stretched bar displays (3840×1100)\n\nOther Important Parts:\n\n• 3× Raspberry Pi 5s\n\n• Multiple OKGEAR/Coolerguys Molex pass-through to dual USB-A 5V adapters — used to power many of the displays directly from the PC power supply.\n\n• Several Elebase USB-A to USB-C adapters — used so USB-C splitters could interface with the Molex-powered USB lines.\n\n• Multiple 1-to-3 USB-C splitters — used to expand display connectivity from the powered USB lines.\n\n• Numerous DbillionDa angled USB-C cables — chosen for their extremely low-profile connectors. I would not recommend these for typical charging use, as they wear out fairly easily but they worked great for this project.\n\n• Sinefine PCIe USB-C power delivery card (4-port) — used to properly power the Raspberry Pi 5s with enough wattage. The Molex USB adapters alone only allowed the Pis to run in low-power mode, which occasionally caused stuttering.\n\n• 3× 52Pi Power HATs — used with the Raspberry Pi 5s to convert the volts/amps from PCIe USB-C P card. Without them, the Pis would still report low-power mode, as they require 5 amps and the USB C card only provides 3 amps.\n\n• 3d Printed Mounts and brackets for holding displays in position, as well as fan/GPU cover (GPU cover between the two horizontal bar displays is made of ABS, a heat resistant plastic, and there's still plenty of airflow going to GPU)","offTopic":true}],"breakdown":[{"sourceKey":"reddit","sourceName":"Reddit","count":4}],"total":4}}