Freestyle Docs

Freestyle / Guides

How to Use Hugging Face with Credential Injection

Call Hugging Face Inference Providers from a Freestyle VM while keeping the real access token at the edge.

Call open models through Hugging Face Inference Providers without storing a real Hugging Face token in the VM. This guide uses the OpenAI-compatible chat-completions endpoint at router.huggingface.co.

Your trusted controller puts the provider key in a Freestyle HTTP egress rule. The VM sends a placeholder; the edge replaces it for the selected endpoint. The real key stays outside the guest filesystem, process environment, and snapshots.

Prepare The Provider Key

Create a fine-grained User Access Token in Hugging Face settings, granting Make calls to Inference Providers. Do not grant Hub repository write access for this inference example.

Choose an available chat model from the Inference Providers model list. The example below uses openai/gpt-oss-120b:fastest; replace it if it is unavailable to your account. The :fastest suffix selects the fastest available provider. Calls are routed and billed through your Hugging Face account, subject to its provider settings and credits.

Set these values on your controller, outside the VM:

export FREESTYLE_API_KEY="your-freestyle-api-key"
export HF_TOKEN="your-provider-credential"
npm install freestyle@latest
npm install --save-dev tsx

The controller needs Node.js 22 or later. The guest example uses Python’s standard library, included in the freestyle/ubuntu base snapshot; no provider SDK is needed.

Create A VM And Grant The Endpoint

Save the following as hugging-face-controller.mts. Keep the real provider key in your controller’s secret store. There is no general public-egress firewall grant: Freestyle creates the named edge route when you add the TLS rule.

import { Freestyle, type CreateTlsRuleOptions } from "freestyle";

function requiredEnv(name: string): string {
  const value = process.env[name]?.trim();
  if (!value) throw new Error(`Missing ${name}`);
  return value;
}

function providerRule(vmId: string, apiKey: string): CreateTlsRuleOptions {
  return {
    action: "allow",
    domain: "router.huggingface.co",
    source: { vmId },
    destination: { public: true },
    match: { method: ["POST"], path: { exact: "/v1/chat/completions" } },
    transform: [{ headers: { "authorization": `Bearer ${apiKey}` } }],
  };
}

const apiKey = requiredEnv("HF_TOKEN");
const freestyle = new Freestyle();
const { vm, vmId } = await freestyle.vms.create({
  snapshotId: "freestyle/ubuntu",
  slug: "hugging-face-client",
  firewall: { rules: [] },
});
console.log({ vmId });
const route = await freestyle.tls.rules.create(providerRule(vmId, apiKey));
console.log({ ruleId: route.id });

The API seals the header value at rest and returns "***" when you read the rule. Any process in this source VM can use the granted endpoint with that credential. Use separate source VMs or VPCs for workloads with different access.

Run The Request Inside The VM

Continue in the same controller file. Write this Python script into the VM:

await vm.fs.writeTextFile(
  "/home/ubuntu/hugging-face-request.py",
  `import json
import ssl
import urllib.error
import urllib.request

payload = {'model': 'openai/gpt-oss-120b:fastest', 'messages': [{'role': 'user', 'content': 'Explain a virtual machine snapshot in one sentence.'}], 'max_tokens': 128, 'stream': False}
request = urllib.request.Request(
    "https://router.huggingface.co/v1/chat/completions",
    data=json.dumps(payload).encode("utf-8"),
    headers={
        "authorization": "Bearer unused-placeholder",
        "content-type": "application/json",
    },
    method="POST",
)
context = ssl.create_default_context(cafile="/etc/ssl/certs/ca-certificates.crt")
try:
    with urllib.request.urlopen(request, context=context, timeout=60) as response:
        result = json.load(response)
except urllib.error.HTTPError as error:
    raise RuntimeError(
        "Hugging Face returned " + str(error.code) + ": " + error.read().decode()
    ) from error

print(result["choices"][0]["message"]["content"])
`,
);

Wait for the hostname mapping and Freestyle CA to reach the guest before the request. This readiness probe opens the provider root without the secret header and without --fail: an HTTP error response still proves the TLS route is reachable. It does not repeat the paid API operation.

let ready = false;
for (let attempt = 0; attempt < 15; attempt++) {
  const check = await vm.exec({
    command: "curl --silent --show-error --output /dev/null --connect-timeout 5 --max-time 10 https://router.huggingface.co/",
    timeoutMs: 15_000,
  });
  if (check.statusCode === 0) {
    ready = true;
    break;
  }
  await new Promise((resolve) => setTimeout(resolve, 2_000));
}
if (!ready) throw new Error("Hugging Face TLS route did not become ready");

const result = await vm.exec({
  command: "python3 /home/ubuntu/hugging-face-request.py",
  timeoutMs: 90_000,
});
if (result.statusCode !== 0) throw new Error(result.stderr ?? "Request failed");
console.log(result.stdout);

Run the assembled file with npx tsx hugging-face-controller.mts. Python explicitly trusts the system CA bundle, including the Freestyle CA. Keep certificate verification enabled. If a request times out, check provider usage before retrying: the operation may already have been accepted.

Choose Models And Providers

Change the guest request’s model to a supported model ID. Hugging Face supports routing suffixes such as :fastest, :cheapest, :preferred, or an explicit provider name. Model and provider availability can change; check the current routing documentation.

The authentication rule does not restrict the model or token budget in the request body. To fix the model from your controller, add a JSON body transform after the header transform:

transform: [
  { headers: { authorization: `Bearer ${apiKey}` } },
  { jsonPatch: [{ op: "add", path: "/model", value: "openai/gpt-oss-120b:fastest" }] },
],

The OpenAI-compatible route covers chat completions. Other inference tasks may use different paths; follow the task’s provider documentation and add only the routes it needs. This rule does not grant private Hub repository access, upload models, or run a model locally inside the VM. Keep repository access on separately scoped credentials and routes.

Check The Access Boundary

RequestExpected behavior
Matching POST from this VM, using the placeholderEdge injects the real key; provider returns a result if the key, quota, and request are valid
Different method or path on the same hostNo credential injection; private endpoints should reject the placeholder
Direct connection to the origin IPBlocked by the absence of public firewall access
Read the TLS ruleReal header value is redacted

The method/path match selects when to inject credentials; it does not deny every other path on the named host. It also does not enforce a per-user usage budget. For authentication failures, check the controller key and exact method/path. For certificate errors, check the guest CA bundle. Provider rate limits and quota errors remain provider errors.

Rotate Credentials And Clean Up

Create a replacement provider credential, load it into the controller, and update the complete rule:

await freestyle.tls.rules.update(
  route.id,
  providerRule(vmId, requiredEnv("HF_TOKEN")),
);

After propagation, verify a new request before revoking the old credential at the provider. A rule read cannot reconstruct the redacted key for an update.

Delete the route and VM when finished, including after a failed example:

await freestyle.tls.rules.delete(route.id);
await vm.delete();

Deleting a route removes this VM’s access through it. Revoke the provider key separately when it should stop working everywhere.

esc