Bringing Your AI Work Home — Setting Up for Local LLMs

In today’s world of AI innovation, having the power to run large language models (LLMs) locally is becoming increasingly valuable. Whether you’re a researcher, developer, or AI enthusiast, bringing your AI work home means taking control of your environment, data privacy, and workflow efficiency without relying on cloud services. With advancements in GPU technology and accessible hardware, setting up an offline LLM workstation is more achievable than ever.

To get started, selecting the right hardware is crucial. A GPU with at least 12 to 16 GB of VRAM, such as the NVIDIA RTX 3060 or RTX 4080 Super, forms the backbone of local LLM work. This ensures you can comfortably run popular models like 7B or even 13B parameter models in quantized formats. Alongside GPU, ample system RAM (48 GB or more) and a fast NVMe SSD help ensure your models load swiftly and inference runs smoothly. Cooling and power considerations, especially in warmer climates, should also factor into your choice to maintain performance and hardware longevity.

Compatibility checks are key before upgrading your PC. Knowing your motherboard’s PCIe slots, available power supply connectors, and physical space within your case avoids costly surprises. For example, HP OEM motherboards typically have PCIe x16 slots suitable for GPUs like the RTX 3060, but verifying available power cables and case clearance is important. While some may wonder about running dual GPUs, for LLM inference, a single, robust GPU is far more practical due to VRAM limitations and software support constraints.

To understand exactly what hardware you have installed, you can use built-in Windows tools such as PowerShell to gather detailed system info. Here are some useful commands:

  • Get-WmiObject Win32_VideoController | Select-Object Name, AdapterRAM — shows your GPU model and memory size.
  • systeminfo | findstr /C:"Total Physical Memory" — reveals your total system RAM.
  • Get-WmiObject Win32_Processor | Select-Object Name, NumberOfCores, NumberOfLogicalProcessors — displays your CPU details.
  • Get-PhysicalDisk | Format-Table FriendlyName, MediaType, Size — lists your storage drives, their type (SSD or HDD), and capacity.

You can run these commands in PowerShell, then copy and paste the results into an online LLM like ChatGPT with a prompt such as:

“Here is my computer hardware info: [paste output]. Please assess my system’s readiness for local LLM work and recommend any upgrades.”

This way, you get tailored advice on how well your current setup fits your AI goals.

Once hardware is in place, your Windows machine (it’s good to also have up-to-date NVIDIA drivers) is ready for software like Ollama, LM Studio, or llama.cpp. Running Ollama with Open WebUI is the fastest way to begin working with LLM models locally (or even offline).

A Gotcha When Setting Up Open WebUI and Ollama on Windows

If you install Ollama natively on Windows while running Open WebUI in a Docker container this may create networking issues because Docker containers on Windows run inside a Linux virtual machine (WSL2), making communication between the container and the Windows host inconsistent. Attempts to connect the two using various IP addresses or hostnames are unreliable.

So what should you do instead? Run Ollama inside a Docker container alongside Open WebUI. This allows the two containers to communicate reliably using Docker’s internal networking.

A final tip– If you previously installed Ollama natively in Windows and downloaded LLM models, (and don’t want to re-download them) you can mount the existing Ollama model directory from Windows into the Docker container using the -v flag. This allows you to reuse previously downloaded models without having to redownload them.

-v C:\Users\YourName\.ollama:/root/.ollama

Running models locally preserves privacy and offers instant access without internet dependency. Bringing AI work home isn’t just about hardware — it’s about creating a streamlined, efficient, and private AI workspace. It’s also really nice that you control all of it.