A man at his desk views a monitor illustrating offline computing and cloud-connected apps.
Parth Shah's week-long experiment at XDA Developers asks a question a lot of Windows users are quietly wondering about. How much of your daily AI work can you move off ChatGPT, Claude and Gemini and onto your own PC? His answer is a hybrid. Local models handle routine, private, bounded jobs well. The cloud still earns its keep when you need current information, connected accounts or project-scale help.

This is one writer's workflow test, not a benchmark. The report gives no hardware specs, quantization levels, context-window settings or measured speeds. Read words like "quick" and "falling apart" as his impressions, not performance guarantees. The useful part is where the lines fell, and why.

A man at his desk views a monitor illustrating offline computing and cloud-connected apps. Setup was the easy part​

Shah says LM Studio and Ollama made getting started painless. He downloaded popular models, including Qwen and Gemma, and had them running quickly. He used them to brainstorm article ideas, build outlines, rewrite clumsy paragraphs, summarize short documents, draft a diet plan and answer basic coding questions. He also mentions the privacy angle: he didn't have to send a prescription to a remote server.

LM Studio's documentation is not in the search results I pulled, but the packet confirms the basics. The app runs on macOS, Windows and Linux. You must download model weights first, after which it can run entirely offline. It can also act as an MCP client and serve local models through OpenAI-compatible endpoints.

On privacy, LM Studio's policy says that when you run models locally, your messages, chat histories and documents aren't transmitted from your system. The exceptions are internet-dependent actions such as searching for or downloading models and checking for app updates. The same policy says cloud features, like web search or cloud models, are separate and apply only if you purchase and use them. Ollama's policy draws a similar line between local use and its cloud-hosted models.

Don't stretch this too far, though. These claims cover the named runners. They don't cover every plug-in, connector or app you bolt onto them.

Summary: For short, private tasks, local AI is easy to set up and good enough that Shah rarely reached for a cloud assistant.

Coding: fine for scripts, shaky for projects​

Shah chose Qwen2.5-Coder 7B running through LM Studio. He reports it handled Python scripts, simple bug fixes, code explanations and small-project edits. It struggled with multi-file work, complicated debugging and larger codebases, which needed repeated prompting and manual intervention. He missed Claude Code's ability to understand a whole project, navigate files and plan changes. That contrast is his experience, not a head-to-head test.

The model card adds useful context. Qwen lists the 7B model at 7.61 billion parameters under an Apache-2.0 license. It shows a full context length of 131,072 tokens. But the card also says the default configuration is set for 32,768 tokens. Longer inputs need YaRN configuration. The card also advises against using the base model for conversation without post-training. Shah doesn't say which variant he downloaded, so I won't guess.

The practical lesson for developers is that an advertised context window isn't what your local setup is actually using. A big window also doesn't guarantee reliable reasoning across many files. LM Studio's local API and MCP support mean you can wire local models into your own workflows. That doesn't mean it matches a cloud coding agent out of the box.

Summary: Local coding help works for bounded tasks. Project-scale work takes more setup, more context management and more supervision.

Research: the stale-knowledge problem​

This was Shah's dealbreaker. Shopping for an OLED monitor, he asked his local model to compare the Samsung Smart Monitor M9 and the Odyssey OLED G8. It didn't know about the latest G8. He then used ChatGPT to compare current models against his needs: a Windows desktop, a MacBook Pro and a Fire TV Stick.

The "latest G8" is real, and it's a good example of why a frozen model gets this wrong. Samsung's announcement describes the Odyssey OLED G8, model G80SH, in 27- and 32-inch sizes. It has an OLED panel with 4K resolution at up to 240Hz, Glare Free technology, and a USB-C port supporting up to 98W. Samsung says it adopts QD-OLED Penta Tandem technology. Samsung's US pricing lists the 27-inch at $1,099.99 and the 32-inch at $1,299.99. Samsung's global release adds that the G80SH supports DisplayPort 2.1 UHBR20, with up to 80Gbps of bandwidth.

One wrinkle is worth knowing. A product-specs roundup noted that Samsung's official materials list the USB-C charging capacity at 98W, while Tom's Hardware reported 96W. Samsung's figure is the one to trust, but check the spec sheet for your region before you buy.

The takeaway isn't "cloud AI is smarter." A local model can reason well over information you paste in. It just can't know about something released after its training data, unless it has a retrieval tool. By the same logic, a cloud model only helps here if web access or another live source is enabled. Hosting alone doesn't make an answer current.

Summary: For anything that changes, such as products, prices, patches or versions, give your model live sources or verify its answers.

Integrations: where the cloud's convenience lives​

Shah says ChatGPT can tie together Notion, Outlook, Google Drive and Asana, and that Gemini works with Google services such as Tasks and Keep, plus apps like WhatsApp and Canva. He says replicating that locally means configuring APIs, connectors and custom workflows, which he didn't want to spend hours on. It is possible, he notes. It just isn't turnkey.

The vendor documentation shows these connections come with conditions:

  • ChatGPT: OpenAI's help page says app availability depends on the app, your plan, region, workspace, role, model and interface. Wait, that detail is from the OpenAI help page in the packet, not this search. In plain terms, apps need account authorization and can be restricted by workspace administrators. OpenAI also says app permissions only control when ChatGPT asks before acting. They don't grant new access.
  • Gemini: Google says Connected Apps vary by Gemini app, device, country and account type. Work and school accounts follow separate rules. For WhatsApp, Google documents sending messages and making calls on Android, so don't read the mention as search access to your message history.
  • Local: LM Studio's MCP support shows local integrations are technically possible. Whether it takes hours depends on the tool.

For IT admins, the point is that connected-app permissions deserve a review before rollout. That applies especially to Outlook and SharePoint-style connections in managed workspaces. Local inference sidesteps that question, but only by giving up the connectors.

Summary: The cloud's edge is turnkey access to your accounts. The cost is permissions, plan limits and data-handling rules you have to understand.

Long documents and the maintenance tax​

Short PDFs went fine for Shah. Larger technical documents meant slower responses, missed details and manual chunking, depending on the model and context settings. He contrasts that with Gemini Notebooks and ChatGPT Projects. OpenAI describes Projects as keeping related chats, files and instructions together, subject to plan and workspace settings. That's a context-management feature, not a promise of unlimited document length.

LM Studio documents an offline "chat with documents" feature using retrieval-augmented generation (RAG). So local document Q&A isn't off the table. Shah's difficulty probably reflects his model, settings and workflow, and that's the honest framing.

He also reports spending a surprising amount of time downloading models, trying quantization levels, adjusting context windows, watching RAM and switching between models for writing, coding and reasoning. He gives no time or memory figures, so treat this as a qualitative warning. Local AI trades a subscription for tinkering time.

A practical split for Windows users​

Based on Shah's experience and the documentation above, a sensible division looks like this:

  1. Keep local: drafting, rewriting, brainstorming, short-document summaries and sensitive personal or work text.
  2. Go cloud or add retrieval: product comparisons, current events, and anything with a recent release date.
  3. Go cloud for connected work: email, calendars, drives and task tools, after checking the permissions you're granting.
  4. Test before trusting: code from a small model, especially across multiple files.
  5. Check the context setting: a model's maximum context isn't necessarily what your runner has loaded.

Is a hybrid setup more work than picking one? Yes. But it matches the evidence, and Shah's own conclusion is that the question isn't local versus cloud. It's knowing when each makes sense.

 

References

  1. I used only local AI for a week and learned what actually needs the cloud XDA 2026-10-09T20:00:17+00:00
  2. Connected apps in ChatGPT | OpenAI Help Center help.openai.com
  3. Use & manage Connected Apps in Gemini - Computer - Gemini Apps Help support.google.com