Your AI models (LLMs) run on your premises.

For some data, an on-premise server is worth comparing with cloud options. We size it to your usage, with network, storage and maintenance clearly scoped.

What you get

  • A GPU server sized to your real usage, installed on your premises
  • Open-weight LLMs with licences checked for the intended use
  • Transcription, speech synthesis and image generation as needed
  • Standard API: your tools connect to it like to a cloud service
  • Network flows, logs, telemetry and support access reviewed

Why a server inside your walls

Medical, legal, industrial and R&D data can benefit from fewer transfers and more control. Updates, telemetry, backups, connected tools and support access still need review. On-site hardware does not remove subprocessors or audit duties by itself.

What we install

We install open-weight models chosen per task, after checking licence and terms: language, transcription, speech or image. An API compatible with your existing client tools can cut integration work, but each model still needs testing.

Sized to your usage, not to marketing

Sizing depends on models, memory, concurrent users, peaks and expected latency. One machine can fit some SME workloads, but not all. Full cost includes energy, cooling, warranty, backups, operations and replacement.

Installed, explained, maintained

Scope can include installation, integration, training and a maintenance agreement. Updates, patches, backups, alerts and remote access each get a clear owner. Continuous monitoring is included only if a service level says so.

Frequently asked questions

What hardware is needed?

A GPU server suited to the models, context, concurrent users and latency target. A tower may fit some uses; others need more memory or several accelerators. We measure before sizing.

Which models can it run?

Recent open-weight models: LLMs for text and documents, transcription, speech synthesis, image generation. We check licences for commercial use before installing anything.

Is it as good as ChatGPT?

For a bounded task, a local model may meet the target, but test it on a representative set. If it misses quality, latency or cost requirements, cloud or hybrid may fit better.

Who handles maintenance?

The contract assigns responsibility to your team, ARCKONE or both. Updates, patches, backups, monitoring and response times need to be explicitly included rather than assumed.

What about the GDPR?

Local deployment can reduce transfers if the whole flow is set up for it. Logs, telemetry, backups, maintenance access, connected tools and security measures still need documenting. Owning hardware does not guarantee GDPR compliance.

We scope it in 20 minutes.

One call is enough to know whether the topic deserves a real project.

Scope an installation