Can we run this ourselves, so none of it ever leaves the building? The organizations I've built private AI for asked exactly that. All of them were small, and all of them handle health information. They did not want a patient note or a member record sitting on someone else's server.

The short answer is yes. I've built exactly this for at least four small organizations, all under 200 staff. I also run it at home, on one consumer graphics card with 12 GB of memory, for transcription and everyday drafting. It is not a science project anymore. It is a few servers, some open software, and a set of decisions you have to make on purpose.

This piece is about those decisions. It is not a setup guide. It covers why you might want your own AI and the three ways to keep AI private (only one of which is an air gap). Then it covers which models are worth your time, what your staff actually see, what it costs in round numbers, and when I'd tell you not to bother. Every number has a source, and where the number is my own arithmetic I say so.

The reason is control, not fear

Let's be honest about the starting point. Privacy is a real worry, but it is not the biggest one. In the Center for Effective Philanthropy's 2025 survey, 67 percent of nonprofit leaders named data security and privacy as a concern about AI, behind inaccurate results at 73 percent.1 Among nonprofits already using AI regularly, only 32 percent named privacy and security as a barrier, per Virtuous's survey of 346 organizations.2

So running your own AI will not fix the thing people worry about most. A private model can still be wrong. What it changes is who decides where your data goes.

That matters more than it used to. The software you already pay for keeps changing its terms, and I wrote about vendors quietly switching on AI training in the tools you already own. AI suites are being bundled as a reason never to leave, which I covered in the lock-in column. And the meter on hosted AI tends to live somewhere other than the demo, as I found in the hidden-meter column.

Meanwhile, your staff are already pasting things into chatbots. Cyberhaven's 2026 report, based on data movements at 222 companies, found that 39.7 percent of AI interactions involved sensitive data.3 That is a vendor measuring its own customers, so treat it as a signal rather than a census. But the direction is not in doubt.

Organizations that host their own AI say the same thing. In Uptime Institute's 2025 survey of data center operators, privacy and security was the top reason for running AI inference on premises, at 68 percent.4 Closer to health care, a 2024 survey of 36 NIH-funded clinical research hubs found 63.9 percent had deployed generative AI on premises.5

The question isn't whether AI is safe. It's whether you get to decide, in writing, where your members' data is allowed to go.

Three ways to keep it private, and only one is an air gap

When people say "private AI," they usually mean one of three different things. They cost very different amounts, so it pays to know which one you actually need.

Option one: a hosted service with a contract that keeps your data out of training

This is the cheapest path and, for many organizations, the right one. The big providers now put it in writing. OpenAI says API data has not been used to train its models by default since March 2023, and it keeps abuse-monitoring logs for up to 30 days.6 Anthropic says it does not train on inputs or outputs from its commercial products by default, and deletes API inputs and outputs within 30 days unless otherwise agreed.7 Microsoft commits that prompts and completions in its Foundry service are not used to train foundation models without permission.8

Read the fine print, though. Amazon Bedrock offers a zero-retention mode, but some of its newest models require a review mode that keeps prompts inside AWS for up to 30 days.9 A contract is only as good as the model you picked under it.

And a word on HIPAA, since that is the reason my clients gave. HIPAA does not require you to run AI on your own servers. It treats a vendor that creates, receives, maintains or transmits protected health information for you as a business associate, which means you need a business associate agreement with it.10 HHS says plainly that a covered entity may use a cloud service to store or process electronic health information, provided it has that agreement in place.10 My clients chose to keep health information off outside systems entirely rather than rely on a vendor's agreement and retention terms. That is a legitimate choice about risk. It is not a legal requirement, and nobody should tell you it is.

Option two: your own servers, with no route to the internet

This is what I've built for my clients. The AI runs on servers in your office or data closet. The servers have no connection to the internet. Staff reach them from their normal computers over the internal network, the same way they reach a shared drive. New models and security patches are downloaded on a separate machine and carried in by hand.

Prompts, documents and answers never leave your network. The trade-off is that someone owns the update routine, and I'll come back to what that costs.

Option three: a true air gap

People use "air-gapped" loosely, so here is the real definition. NIST defines an air gap as an interface where systems are not physically connected, and any logical connection is manual and under human control.11 By that standard, option two is not an air gap, because the servers talk to computers that also talk to the internet. It is isolated, which is a strong control, but it is a different thing.

A true air gap means the people using the AI sit on a separate network too. It exists, and the tools support it: NVIDIA documents moving model files to a disconnected system on approved media.12 But every model, patch and container image then crosses the gap by hand. Very few associations have data that justifies that.

One trap to watch for. Some chat tools have an "offline mode" setting. Open WebUI's, for example, turns off update checks and model downloads, but it does not block connections to outside AI services.13 An app setting is not an air gap. Only the network is.

The open models are good enough for the work you actually do

"Open-weight" models are ones you can download and run on your own hardware. They used to be a clear step down. Now they are close enough for most association work, as long as you are honest about the gap.

Here is the gap. On Artificial Analysis's leaderboard in late September 2026, the best closed model scored 58 and the best open-weight model scored 46.14 Epoch AI puts the delay another way: since January 2026, the most capable open models have trailed the frontier by an average of four months.15

Four months behind the best is still very good at drafting a renewal letter, summarizing a board packet or answering a question from your own policies. Those benchmarks mostly test hard reasoning. For routine writing and summarizing, I think the difference is small. Where it shows up is on complex analysis and long multi-step tasks, so plan those around a hosted model or a human.

What fits on what

You do not need a data center. Google says its Gemma 4 12B model is small enough to run on a laptop with 16 GB of memory.16 OpenAI's gpt-oss-20b runs within 16 GB of memory, and its larger gpt-oss-120b fits on a single 80 GB graphics card.17 Both gpt-oss models and all of Gemma 4 use the Apache 2.0 license, which lets you use them commercially without a per-user fee.18

My shortlist for a small organization looks like this:

For searching your own documents you also need a smaller "embedding" model. IBM's granite-embedding and Google's EmbeddingGemma are both built for exactly this, and the Google one is designed to run on a laptop.21

Where a model comes from matters now

Notice who is missing from that list. Meta's newest open models are still Llama 4, from April 2025.22 The strongest open models today come from Chinese labs, and that has consequences. In January 2026, Texas extended its list of technologies banned on state devices to Chinese AI companies including Alibaba, maker of the popular Qwen models, and BAAI, maker of the widely used BGE embedding models.23

That rule covers Texas state agencies, not you. But if you have state contracts, or a board that will ask, it is easier to start with a US or European model than to explain later. And hosting a model yourself removes the data risk, not the behavior risk. When NIST's AI safety center tested DeepSeek R1 in 2025, it answered 94 percent of overtly malicious requests under a common jailbreak, against 8 percent for US reference models.24

What your staff actually see

Here's the part people don't expect. For staff, private AI looks almost exactly like the public kind. They open a browser, sign in with the same Microsoft account they use for email, and get a chat window. They can upload a document, ask about it, and get an answer.

Behind that window sit three pieces. There's an engine that runs the model (Ollama for simplicity, vLLM when many people use it at once). There's a chat front end such as Open WebUI or LibreChat, both of which support single sign-on through Microsoft Entra ID.25 And there's the part that searches your own documents.

Check the licenses. Open WebUI is no longer open source in the formal sense. Since April 2025, its license requires the Open WebUI branding to stay visible unless a deployment has 50 or fewer users in a 30-day period or buys an enterprise license.26 Any number of internal users is fine if you keep the logo. LibreChat uses the MIT license and has no such condition.25

The hard part is permissions, not models

The one tricky piece is document search. When you point an AI at your shared drives, it can find anything it has been given, including files a given person was never supposed to see. OWASP's 2025 list of AI application risks warns that misaligned access controls can expose sensitive data, and that one group's content can be retrieved for another group's question.27

Microsoft's Copilot handles this by only showing data the user can already open, and even Microsoft tells customers to clean up their SharePoint permissions first.28 A private system has to do the same thing on purpose. In practice that means starting with one well-governed set of documents, such as your policies, your standards or your knowledge base, rather than everything.

What it costs, in round numbers

Here is where I have to be straight with you. On cost alone, running your own AI does not beat renting it at association size. The case is control. Let me show the pieces so you can judge for yourself.

Hardware

Graphics card prices have been volatile all year. NVIDIA's store listed the RTX PRO 6000 at $16,000 in August 2026, up from $13,250 in June.33 Get a current quote before you budget.

Power and depreciation

Electricity is the smallest line. By my arithmetic, a DGX Spark running flat out around the clock would use about $290 of power a year at the 13.85 cent US commercial average. The two cards in the department server would use about $1,460 a year under the same assumptions, before cooling. Real use is far lower.

Depreciation is bigger. Amazon now depreciates many of its servers over five years.34 Spread over five years, that $68,011 server costs about $13,600 a year before anyone touches it.

People

This is the line that decides it. Hardware you buy once. Care and feeding you pay for every month.

A network and systems administrator in the Washington, DC area earned a median $125,430 in May 2025, per the Bureau of Labor Statistics.35 With benefits, which BLS puts at 31.5 percent of compensation for professional roles, that is about $183,000 a year fully loaded. You will not hire one for this. The realistic model is your existing IT person plus an outside partner.

The work is steady, not dramatic. vLLM, the most common engine for serving many users, published 59 security advisories in 2026 through late September, against 27 in all of 2025.36 Joint security guidance led by the NSA, with CISA, the FBI and allied agencies, is written for exactly these on-premises deployments. It says to patch the environment, and to run a full evaluation of accuracy, performance and security whenever you update the model, before putting it back in service.37 In an isolated setup, each of those updates is carried in by hand.

On the builds I've done, the work runs in stages. First comes a network assessment, which takes one to two weeks. Then a written proposal, a review with your team, and the build itself. How long the build takes depends on the infrastructure you already have: weeks for some organizations, a few months for others. Most are up and running in four to six weeks. After that, the ongoing work is a monthly patch window and a test before every model change.

What renting the same thing costs

Now compare. Take a heavy-use assumption of my own: 200 people each sending 50 requests a working day, at about 3,000 tokens (the chunks AI models count text in) per request. At Together AI's published prices, that would cost about $3,000 a year on gpt-oss-120b and about $8,200 a year on Llama 3.3 70B.38 On Anthropic's Claude Sonnet 5.5, a top closed model, it would be about $47,500 a year.39

You will see studies that say owning is far cheaper. Read who paid for them. A Dell-commissioned study found on-premises AI up to 4.1 times more cost-effective than a leading API, but its smallest scenario was 5,000 users on eight H100 cards.40 A Lenovo paper claims up to a 17-fold advantage, for high-utilization workloads over five years.31 Those are real results for big, busy systems. A 60-person association is neither.

Deloitte's interviewees offered a better rule of thumb: start in the cloud, and consider owning hardware once the cloud bill for a workload reaches 60 to 70 percent of the cost of buying the systems.41 For most small organizations, that day never comes on cost. You own the servers because of what they do not send anywhere.

What people actually use it for

The uses are ordinary, and that is the point. In CEP's survey, almost two-thirds of nonprofits and foundations used AI, mainly for internal productivity: drafting emails, policies and procedures, and summarizing meetings.1 Virtuous found the most common uses were donor communications at 62 percent, marketing and social media at 60 percent, and data analysis and reporting at 42 percent.2

At the organizations where I've built private AI, staff lean on it for four things:

Every one of those is text that already lives inside the building. That is why a private model fits them so well.

When I'd tell you not to bother

I'd rather lose the project than sell you a server you don't need. Here is when I'd push back.

When the hardware would sit idle. Among enterprises that run their own graphics cards, 69 percent report using half their capacity or less, according to a July 2026 survey of 155 of them.42 Big companies with dedicated teams struggle to keep these machines busy. A small office will too.

When a contract would do. If your worry is vendor training on your data, option one probably covers you for a fraction of the cost. Governments reach the same conclusion. France dropped its in-house Albert assistant after pilots in 48 service centers produced wrong answers, and its replacement runs on a certified French cloud provider rather than government-owned servers.43

When nobody will own it. In CEP's survey, 58 percent of nonprofit leaders cited a lack of staff expertise.1 A private AI with no owner becomes an unpatched server with a lot of sensitive data on it. That is worse than where you started.

This isn't an argument against hosted AI. For many organizations, a well-negotiated contract is the right call. But if you handle health information, or your board has drawn a line, it's worth knowing the other door is open and not as expensive to walk through as people assume.

Where to start

Start by writing one sentence: which data must never leave the building? If you can't name any, you probably need a better contract, not a server. If you can, test a small model on a single box with the people who handle that data, on one set of documents, before you buy anything larger.

Either way, the decision is yours to make on purpose, and not by default in someone else's terms of service.