AI for Enterprise

Enterprise Small Language Model Vendors in 2026: Who Builds, Owns, and Self-Hosts the Model Layer

Written by
Dr. Anushtha Singh
Created On
06 Oct, 2026

Table of Contents

Don’t miss what’s next in AI.

Subscribe for product updates, experiments, & success stories from the NuPlay team.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

_Disclosure: NuPlay AI builds a competing platform. This comparison is limited to the public sources cited below; it does not establish private deployment configurations, self-serve trial results, or contractual terms. _Source

To evaluate enterprise small language model vendors, look beyond parameter count and verify whether a vendor builds a named model, who holds and can deploy the weights, where inference runs, and whether any external model provider remains in the production path.

Writer’s September 2025 release shows why this distinction matters. Its Palmyra-mini family was released as open models, with variants listed at 1.5B and 1.7B parameters. An enterprise small language model can therefore be an openly available model that a customer places independently, rather than only a vendor-hosted service. Writer’s release.

A small language model (SLM) is a language model with far fewer parameters than a frontier large language model, built or fine-tuned for a narrower set of tasks and small enough to run inside infrastructure a customer controls. Gartner predicts that by 2027, organizations will use small, task-specific AI models at least three times more than general-purpose large language models, and Gartner VP Analyst Sumit Agarwal says these models provide “quicker responses and use less computational power.” Gartner, April 9, 2025

NVIDIA Research makes a related argument for agentic systems in a 2025 position paper, which proposes small models for narrow, repetitive invocations. It is a position piece rather than a new benchmark, and a preprint.

What is a small language model vendor?

For this comparison, a small language model vendor is a provider that can document one or more of three model-layer positions:

  1. It builds a named model family.
  2. It provides or licenses usable model weights.
  3. It can run inference inside a customer-controlled virtual private cloud (VPC), on-premises environment, or air-gapped boundary.

Whether external model-provider calls remain in the production path is a fourth question that applies to every vendor on the list, and the verification test below covers it.

These are separate procurement questions. Deployment location, model intellectual-property ownership, possession of weights, and data-control terms are not interchangeable. A customer VPC deployment does not by itself mean the customer owns the weights. An open model does not by itself prove that the vendor operates it for the customer.

NuPlay AI’s Astra page and Writer’s Palmyra-mini release illustrate the distinction between vendor-operated infrastructure and openly available model weights. Astra is NuPlay AI’s family of enterprise AI models; Astra SLM is its smaller, specialized member, built for high-volume, narrow-scope voice and chat work. For fundamentals, see NuPlay AI’s explainer, What is a small language model (SLM)?

Compare public evidence, not marketing categories

How do enterprise SLM vendors compare on public evidence?

The following comparison separates what each vendor publicly documents from what remains unverified in public material. It is not a ranking or a substitute for contract and architecture review.

Vendor Publicly documented model-layer position What the buyer should not infer Evidence status
NuPlay AI Astra SLM is a 9B in-house specialist model for high-volume voice and chat, deployable inside a customer’s VPC. NuPlay AI states that Astra is self-hosted with zero third-party model dependencies. Astra Do not say that a customer owns Astra weights. The public page supports vendor ownership and customer-infrastructure deployment, not a transfer of model intellectual property. Vendor claim
Writer Writer describes Palmyra-mini as a family of completely open models designed to run privately on a customer’s infrastructure or devices. Writer source “Open” does not settle enterprise support, indemnity, deployment operations, or contractual rights. Confirm the applicable model license and support terms. Documented
Decagon Decagon says its agents use both frontier-provider models and models developed by Decagon Labs. In one customer-VPC deployment, it documented private connectivity to frontier providers alongside self-hosted inference for Decagon-developed models. Decagon source, July 9, 2026 Do not call this a no-dependency stack. Its own account describes a mixed model layer that can retain private connectivity to external providers. Documented
Palantir Palantir AIP documents self-hosting open-source or custom models on an organization’s own infrastructure through a compute module and registered model workflow. Palantir documentation This documents a bring-your-own-model deployment path, not that Palantir builds a named enterprise SLM whose weights customers receive. Documented
C3 AI C3 AI documents deployment of open-source and fine-tuned models within a customer’s on-premises or cloud environment, including models selected from an external model hub or a customer filesystem. C3 AI source Do not treat platform support for a model as C3 AI ownership of that model’s weights. Documented
Aisera Aisera publicly describes task-specific SLMs alongside proprietary and third-party foundation-model options. Aisera source The reviewed public page does not establish whether customers receive weights for task-specific SLMs or whether those SLMs can run in a customer VPC or on-premises. Not publicly verifiable

Evidence status uses three levels. Documented: stated in vendor documentation, a legal or compliance document, or a publicly checkable artifact such as a model card. Vendor claim: stated only in marketing pages, blog posts or launch posts. Not publicly verifiable: public material does not settle the question; ask in a demo or contract review. None of these levels means a customer’s contract, production architecture, or security controls have been independently verified.

Public product pages also cannot prove that a vendor transfers model intellectual property, grants ownership of proprietary weights, supports air-gapped deployment in every configuration, or eliminates every external dependency. Treat the table as shortlist evidence, then confirm the relevant rights and architecture in writing.

Use a four-question verification test before shortlisting

How do you verify an SLM vendor's claims before shortlisting?

A vendor’s answer should cover the whole inference path, not only the hosting label. Decagon’s deployment account is useful here because it distinguishes a dedicated VPC in the vendor’s cloud from a full-stack deployment in the customer’s cloud environment. Decagon’s deployment description.

1. Who built the model?

Vendor-built models are often fine-tuned from an open-weight base. That is acceptable, but ask for the base model and its license terms.

Request the model family name, creator, versioning policy, and evidence of whether the model is:

  • Built by the vendor
  • Open-weight or openly available
  • Licensed from another provider
  • A combination of those sources

A platform that can deploy a model is not necessarily the organization that created it. Ask which model is used for each task, including classification, extraction, generation, routing, and fallback behavior.

2. Who controls the weights?

Ask whether your organization can:

  • Receive and retain a copy of the weights
  • Modify or fine-tune the weights
  • Independently serve the model
  • Keep access after the vendor relationship ends
  • Control model updates and rollback

Access to an application programming interface (API) is not possession of weights. Vendor ownership, customer custody, and open availability are different arrangements and should appear separately in the contract.

3. Where does each inference call run?

Require an architecture diagram that shows:

  • The VPC or on-premises boundary
  • The location of model-serving infrastructure
  • Any vendor-managed environment
  • Any external endpoint
  • Data movement between workflow components

“Private deployment” can describe several arrangements. A dedicated environment in a vendor’s cloud is not the same as a model stack running inside your cloud account. An on-premises option may also differ from an air-gapped configuration.

Also ask who holds the encryption keys, where logs are stored and for how long, and whether your conversations are used to train the vendor’s models. Self-hosting puts these controls within reach. It does not set them.

4. Which dependencies remain?

Ask for a complete inference-path inventory covering:

  • Model APIs
  • Routing layers
  • Telemetry
  • Safety services
  • Speech services
  • Authentication services
  • Update channels
  • External connectivity

A vendor can self-host one model while calling an external provider for another. That may be acceptable for your use case, but it should be an explicit design choice rather than an assumption based on the word “self-hosted.”

A “single-tenant” answer only addresses tenancy. It does not establish customer-VPC deployment, air-gap operation, or ownership of weights.

Worked example: a private claim-intake workflow

What does a private claim-intake workflow look like in practice?

Hypothetical scenario: A US insurer wants to move high-volume claim-intake classification and document extraction into agent mode. It requires a customer-controlled AWS VPC, no public inference calls, and a documented rollback path for model updates.

Apply the four-question test to the available deployment routes.

The following table is a decision aid, not a recommendation for a specific vendor.

Deployment route Likely fit for the insurer Questions that remain
Dedicated software as a service (SaaS) only Does not meet the stated placement requirement if inference remains outside the insurer’s environment. Whether any private alternative exists and which calls leave the vendor environment.
Vendor’s proprietary model inside the insurer’s VPC May meet the runtime requirement. Whether the insurer receives or controls the weights, who operates updates, and how rollback works.
Open-model deployment operated by the insurer May provide greater control over model placement and weights. Who handles patching, evaluation, capacity planning, monitoring, and incident response.
Mixed architecture with private external-provider connectivity May be acceptable if the insurer permits external provider dependency through approved private connections. Which requests use external providers, what data crosses the connection, and how provider changes are governed.

The insurer should not use model size as the success criterion. It should first confirm that the proposed architecture meets the placement and dependency requirements.

A practical proof of concept could replay 50 to 100 representative, de-identified workflow cases and 15 to 20 failure or edge cases. The team should require network logs showing whether any prohibited external model endpoint was contacted.

The evaluation should also check:

  • The agreed workflow outcome
  • A traceable model version
  • Evidence that the model runs inside the required VPC
  • A tested rollback
  • A documented response when a dependency becomes unavailable

This approach tests the operating arrangement the insurer is buying, not only the model’s published description.

Select the control level your workflow needs

Which control level does your workflow need?

Self-hosting is not automatically the right choice. It can increase control over placement and data movement, but it can also move responsibility for infrastructure, capacity, updates, monitoring, and incident response to your organization.

The right shortlist depends on the workflow’s requirements:

  • Choose an open-model route when your team needs direct control of weights and can operate the serving layer.
  • Choose vendor-operated model infrastructure in your VPC when you need private placement but want the vendor to operate much of the model layer.
  • Consider a mixed architecture when external model connectivity is allowed and the added capability is worth the dependency.
  • Reject a deployment label that does not come with an architecture diagram and an inference-path inventory.

For enterprises that need the vendor to own and operate the model layer inside their environment, NuPlay AI describes its operating model this way:

NuPlay AI runs enterprise workflows in production and improves them after every run through NuLoop, with agents, the systems they operate, and the context they draw on under one platform.

In that model, NuLoop moves through Report, Diagnose, Propose, Try, and Ship. Changes are tested against real past runs, can be rolled back, and require human approval before promotion. The relevant question is not whether a vendor uses a particular model label. It is whether the vendor can document who controls the model layer, where it runs, which dependencies remain, and how governed changes reach production.

Talk to the team

Conversational AI for Sales and Support teams

Talk to our team to see how to see how Nuplay AI powers smarter engagement.

Let’s Talk

Ready to see what agentic AI can do for your business?

Book a quick demo with our team to explore how NuPlay AI can automate and scale your workflows

Let’s Talk
Does deploying a small language model in a VPC mean my company owns the weights?
No. Deployment rights and weight ownership are separate contractual questions. A vendor may deploy its own model inside your VPC while retaining ownership of the weights. Ask whether you can receive, retain, modify, independently serve, and continue using the weights after the agreement ends.
Can an enterprise self-host a model without building its own?
Yes. Palantir and C3 AI publicly document routes for hosting open-source or custom models on an organization’s own infrastructure. Neither cited page is specific to small language models, so confirm support for the model size and family you intend to run.
How can I verify whether a vendor still calls an external model provider?
Ask for an inference-path diagram, an egress inventory, and network-log evidence from the proof of concept. The review should cover model calls, routing, telemetry, safety services, speech services, and update channels. A vendor’s statement that its deployment is private does not, by itself, identify every dependency.
Which small language model vendors publicly document customer-VPC deployment?
Use the comparison table and retain its evidence-status labels rather than making a broad unsupported claim. [NuPlay’s Astra page](https://www.nuplay.ai/astra) describes Astra SLM as deployable inside a customer’s VPC. Decagon’s deployment account describes a full-stack deployment in a customer’s cloud environment, while Palantir and C3 AI document ways to run open-source or custom models on an organization’s infrastructure. These sources do not establish identical deployment options, model ownership, or contractual rights.
Related

Related Blogs

Explore All
<---NEW-FAQ--->