_Disclosure: NuPlay AI builds a competing platform. This comparison is limited to the public sources cited below; it does not establish private deployment configurations, self-serve trial results, or contractual terms. _Source
To evaluate enterprise small language model vendors, look beyond parameter count and verify whether a vendor builds a named model, who holds and can deploy the weights, where inference runs, and whether any external model provider remains in the production path.
Writer’s September 2025 release shows why this distinction matters. Its Palmyra-mini family was released as open models, with variants listed at 1.5B and 1.7B parameters. An enterprise small language model can therefore be an openly available model that a customer places independently, rather than only a vendor-hosted service. Writer’s release.
A small language model (SLM) is a language model with far fewer parameters than a frontier large language model, built or fine-tuned for a narrower set of tasks and small enough to run inside infrastructure a customer controls. Gartner predicts that by 2027, organizations will use small, task-specific AI models at least three times more than general-purpose large language models, and Gartner VP Analyst Sumit Agarwal says these models provide “quicker responses and use less computational power.” Gartner, April 9, 2025
NVIDIA Research makes a related argument for agentic systems in a 2025 position paper, which proposes small models for narrow, repetitive invocations. It is a position piece rather than a new benchmark, and a preprint.
What is a small language model vendor?
For this comparison, a small language model vendor is a provider that can document one or more of three model-layer positions:
- It builds a named model family.
- It provides or licenses usable model weights.
- It can run inference inside a customer-controlled virtual private cloud (VPC), on-premises environment, or air-gapped boundary.
Whether external model-provider calls remain in the production path is a fourth question that applies to every vendor on the list, and the verification test below covers it.
These are separate procurement questions. Deployment location, model intellectual-property ownership, possession of weights, and data-control terms are not interchangeable. A customer VPC deployment does not by itself mean the customer owns the weights. An open model does not by itself prove that the vendor operates it for the customer.
NuPlay AI’s Astra page and Writer’s Palmyra-mini release illustrate the distinction between vendor-operated infrastructure and openly available model weights. Astra is NuPlay AI’s family of enterprise AI models; Astra SLM is its smaller, specialized member, built for high-volume, narrow-scope voice and chat work. For fundamentals, see NuPlay AI’s explainer, What is a small language model (SLM)?
Compare public evidence, not marketing categories
How do enterprise SLM vendors compare on public evidence?
The following comparison separates what each vendor publicly documents from what remains unverified in public material. It is not a ranking or a substitute for contract and architecture review.
Evidence status uses three levels. Documented: stated in vendor documentation, a legal or compliance document, or a publicly checkable artifact such as a model card. Vendor claim: stated only in marketing pages, blog posts or launch posts. Not publicly verifiable: public material does not settle the question; ask in a demo or contract review. None of these levels means a customer’s contract, production architecture, or security controls have been independently verified.
Public product pages also cannot prove that a vendor transfers model intellectual property, grants ownership of proprietary weights, supports air-gapped deployment in every configuration, or eliminates every external dependency. Treat the table as shortlist evidence, then confirm the relevant rights and architecture in writing.
Use a four-question verification test before shortlisting
How do you verify an SLM vendor's claims before shortlisting?
A vendor’s answer should cover the whole inference path, not only the hosting label. Decagon’s deployment account is useful here because it distinguishes a dedicated VPC in the vendor’s cloud from a full-stack deployment in the customer’s cloud environment. Decagon’s deployment description.
1. Who built the model?
Vendor-built models are often fine-tuned from an open-weight base. That is acceptable, but ask for the base model and its license terms.
Request the model family name, creator, versioning policy, and evidence of whether the model is:
- Built by the vendor
- Open-weight or openly available
- Licensed from another provider
- A combination of those sources
A platform that can deploy a model is not necessarily the organization that created it. Ask which model is used for each task, including classification, extraction, generation, routing, and fallback behavior.
2. Who controls the weights?
Ask whether your organization can:
- Receive and retain a copy of the weights
- Modify or fine-tune the weights
- Independently serve the model
- Keep access after the vendor relationship ends
- Control model updates and rollback
Access to an application programming interface (API) is not possession of weights. Vendor ownership, customer custody, and open availability are different arrangements and should appear separately in the contract.
3. Where does each inference call run?
Require an architecture diagram that shows:
- The VPC or on-premises boundary
- The location of model-serving infrastructure
- Any vendor-managed environment
- Any external endpoint
- Data movement between workflow components
“Private deployment” can describe several arrangements. A dedicated environment in a vendor’s cloud is not the same as a model stack running inside your cloud account. An on-premises option may also differ from an air-gapped configuration.
Also ask who holds the encryption keys, where logs are stored and for how long, and whether your conversations are used to train the vendor’s models. Self-hosting puts these controls within reach. It does not set them.
4. Which dependencies remain?
Ask for a complete inference-path inventory covering:
- Model APIs
- Routing layers
- Telemetry
- Safety services
- Speech services
- Authentication services
- Update channels
- External connectivity
A vendor can self-host one model while calling an external provider for another. That may be acceptable for your use case, but it should be an explicit design choice rather than an assumption based on the word “self-hosted.”
A “single-tenant” answer only addresses tenancy. It does not establish customer-VPC deployment, air-gap operation, or ownership of weights.
Worked example: a private claim-intake workflow
What does a private claim-intake workflow look like in practice?
Hypothetical scenario: A US insurer wants to move high-volume claim-intake classification and document extraction into agent mode. It requires a customer-controlled AWS VPC, no public inference calls, and a documented rollback path for model updates.
Apply the four-question test to the available deployment routes.
The following table is a decision aid, not a recommendation for a specific vendor.
The insurer should not use model size as the success criterion. It should first confirm that the proposed architecture meets the placement and dependency requirements.
A practical proof of concept could replay 50 to 100 representative, de-identified workflow cases and 15 to 20 failure or edge cases. The team should require network logs showing whether any prohibited external model endpoint was contacted.
The evaluation should also check:
- The agreed workflow outcome
- A traceable model version
- Evidence that the model runs inside the required VPC
- A tested rollback
- A documented response when a dependency becomes unavailable
This approach tests the operating arrangement the insurer is buying, not only the model’s published description.
Select the control level your workflow needs
Which control level does your workflow need?
Self-hosting is not automatically the right choice. It can increase control over placement and data movement, but it can also move responsibility for infrastructure, capacity, updates, monitoring, and incident response to your organization.
The right shortlist depends on the workflow’s requirements:
- Choose an open-model route when your team needs direct control of weights and can operate the serving layer.
- Choose vendor-operated model infrastructure in your VPC when you need private placement but want the vendor to operate much of the model layer.
- Consider a mixed architecture when external model connectivity is allowed and the added capability is worth the dependency.
- Reject a deployment label that does not come with an architecture diagram and an inference-path inventory.
For enterprises that need the vendor to own and operate the model layer inside their environment, NuPlay AI describes its operating model this way:
NuPlay AI runs enterprise workflows in production and improves them after every run through NuLoop, with agents, the systems they operate, and the context they draw on under one platform.
In that model, NuLoop moves through Report, Diagnose, Propose, Try, and Ship. Changes are tested against real past runs, can be rolled back, and require human approval before promotion. The relevant question is not whether a vendor uses a particular model label. It is whether the vendor can document who controls the model layer, where it runs, which dependencies remain, and how governed changes reach production.
.gif)






