Cloud 12 — a SageMaker training job for f4-lora, and the endpoint that bills whether or not it is asked anything

In Java, "deploy the model" and "run the job that trained it" are the same kind of resource — a process that exists while it runs, and stops costing anything the moment it exits. SageMaker's training job is exactly that: it runs, produces a model artifact in S3, and its billing ends the second it reaches Completed. Its real-time endpoint is not — it is a server that keeps running, and keeps billing per hour, until something explicitly deletes it, whether the last request arrived one second ago or twelve hours ago. This page runs f4-lora's training on ml.g4dn.xlarge, deploys the adapter to an endpoint, invokes it exactly once to prove it answers, and deletes the endpoint the same session — because the one thing worse than not shipping this card is shipping it and going to bed with the endpoint still up.

~2 h hands-onml.g4dn.xlarge ~$0.74/h (approx.)delete same day cloud-12

What this creates

Three things, in the order they exist: a training job (sagemaker.pytorch.PyTorch estimator, launched via estimator.fit()) that runs f4-lora's LoRA fine-tuning script on one ml.g4dn.xlarge instance and exits; a model artifact (model.tar.gz) the job writes to S3 when it finishes; and a real-time endpoint (estimator.deploy()) that loads that artifact onto its own ml.g4dn.xlarge instance and answers inference requests over HTTPS. The first two stop costing anything once the job says Completed — SageMaker training bills per second the job actually runs, nothing after. The endpoint is the opposite: it is a server, not a job, and it bills for every hour it exists, occupied or not. The figure below (~$0.74/hour, on-demand, approx.) is representative of published ml.g4dn.xlarge SageMaker pricing — verify on the pricing page today before you rely on it for real money. Four steps below move the scene from "nothing exists" through "invoked once" to "left overnight" so the idle-endpoint number is something you compute once and never forget.

Preconditions

Do it

1. The estimator. Pin the framework version; ml.g4dn.xlarge is the smallest GPU instance SageMaker offers that fits a 7B LoRA fine-tune in fp16/bf16 with a 4-bit base:

# train_lora.py driver — run from a notebook or a local Python shell with the sagemaker SDK
# pip install "sagemaker>=2.220,<3" (pinned; the SDK's estimator API has changed across majors)
import sagemaker
from sagemaker.pytorch import PyTorch

role = "arn:aws:iam::<account-id>:role/<your-sagemaker-execution-role>"
session = sagemaker.Session()

estimator = PyTorch(
    entry_point="train.py",              # f4-lora's own PEFT LoRA training script, unchanged
    source_dir="./f4-lora-src",           # train.py + requirements.txt (peft, transformers, bitsandbytes, pinned)
    role=role,
    framework_version="2.3",              # pinned PyTorch container version
    py_version="py311",
    instance_type="ml.g4dn.xlarge",       # 1x T4 16GB — the smallest GPU instance that fits this job
    instance_count=1,
    hyperparameters={
        "r": 16, "alpha": 32, "target_modules": "q_proj,v_proj",
        "epochs": 3,
    },
    output_path="s3://<your-bucket>/f4-lora/artifacts",
)

2. The job. fit() blocks until the job reaches a terminal state; this is the only step that bills by the second, and only while it runs:

estimator.fit({"training": "s3://<your-bucket>/f4-lora/data/train"})
# ... CloudWatch logs stream here ...
# 2024-xx-xx INFO training job completed
# Training seconds: 5,412   Billable seconds: 5,412   -> ~1.5 h x ~$0.74/h (approx.)

3. The endpoint. This is the resource that behaves differently from the job above — it does not exit when it is done, because a real-time endpoint is never "done":

predictor = estimator.deploy(
    initial_instance_count=1,
    instance_type="ml.g4dn.xlarge",
    endpoint_name="f4-lora-endpoint",
)
# -------------!  <- SageMaker's own progress dots; this line alone means "billing has started"

4. One invocation. Prove it answers, once — this project needs exactly one real request, not a load test:

result = predictor.predict({"inputs": "Summarize: " + open("sample.txt").read()})
print(result)
# {"generated_text": "..."}  -- one request; the endpoint keeps billing after this returns

5. delete_endpoint — the same session, before you do anything else.

predictor.delete_endpoint()
# or, from the CLI if the notebook session is gone:
aws sagemaker delete-endpoint --endpoint-name f4-lora-endpoint
aws sagemaker delete-endpoint-config --endpoint-config-name f4-lora-endpoint
aws sagemaker delete-model --model-name <model name from `aws sagemaker list-models`>

Verify

Teardown

aws sagemaker delete-endpoint --endpoint-name f4-lora-endpoint 2>/dev/null
aws sagemaker delete-endpoint-config --endpoint-config-name f4-lora-endpoint 2>/dev/null
aws sagemaker list-models --query 'Models[].ModelName' --output text | tr '\t' '\n' | grep f4-lora \
  | xargs -I{} aws sagemaker delete-model --model-name {}
# stop any notebook instance you used to drive this, if one was left running
aws sagemaker list-notebook-instances --query 'NotebookInstances[?NotebookInstanceStatus==`InService`].NotebookInstanceName' --output text \
  | tr '\t' '\n' | xargs -I{} aws sagemaker stop-notebook-instance --notebook-instance-name {}

aws sagemaker list-endpoints --query 'Endpoints[].EndpointName'
# []

What remains and what it costs, approx., verify on the pricing page today: the model.tar.gz artifact and any training data in S3 (a few cents per GB-month, the same "keep 5" instinct as cloud-05's ECR lifecycle policy), and nothing else — the training job cannot be "deleted" because it already stopped billing the moment it completed; only the endpoint, its config and the model registration needed an explicit delete, and all three are gone above.

This is self-attestation — the site cannot see your AWS account, so checking the box and pressing the button is you telling The Path the endpoint answered one real request and is now gone.

Takeaways: a training job and a real-time endpoint are billed on completely different clocks, and treating them the same is how a ~$1 fine-tune turns into a ~$10 surprise (the scene's step 4: $1.10 of training, $8.84 of endpoint nobody was using). The job's clock stops itself — Completed is also "billing over". The endpoint's clock does not know the difference between "answering real traffic" and "sitting there since last night" — only delete_endpoint stops it, which is exactly why this page's mark-done will not file until you have run it and proven list-endpoints empty.