In Java, "deploy the model" and "run the job that trained it" are the same kind
of resource — a process that exists while it runs, and stops costing anything the moment it
exits. SageMaker's training job is exactly that: it runs, produces a model artifact in
S3, and its billing ends the second it reaches Completed. Its
real-time endpoint is not — it is a server that keeps running, and keeps billing
per hour, until something explicitly deletes it, whether the last request arrived one second
ago or twelve hours ago. This page runs f4-lora's training on ml.g4dn.xlarge,
deploys the adapter to an endpoint, invokes it exactly once to prove it answers, and deletes
the endpoint the same session — because the one thing worse than not shipping this card is
shipping it and going to bed with the endpoint still up.
Three things, in the order they exist: a training job
(sagemaker.pytorch.PyTorch estimator, launched via estimator.fit())
that runs f4-lora's LoRA fine-tuning script on one ml.g4dn.xlarge instance and
exits; a model artifact (model.tar.gz) the job writes to S3 when it
finishes; and a real-time endpoint (estimator.deploy()) that loads that
artifact onto its own ml.g4dn.xlarge instance and answers inference requests over
HTTPS. The first two stop costing anything once the job says Completed — SageMaker
training bills per second the job actually runs, nothing after. The endpoint is the opposite:
it is a server, not a job, and it bills for every hour it exists, occupied or not. The figure
below (~$0.74/hour, on-demand, approx.) is representative of published
ml.g4dn.xlarge SageMaker pricing — verify on the pricing page today before you
rely on it for real money. Four steps below move the scene from "nothing exists" through
"invoked once" to "left overnight" so the idle-endpoint number is something you compute once
and never forget.
train.py this
page hands to SageMaker should already have run at least one short pass on your own GPU or
CPU, so the first time it runs on a billed instance is not also the first time it runs at
all.aws s3 cp or sagemaker.Session().upload_data(...)), with its
S3 URI noted for step 1.AmazonSageMakerFullAccess-equivalent trust policy for
sagemaker.amazonaws.com. Note its ARN for step 1.1. The estimator. Pin the framework version; ml.g4dn.xlarge is the
smallest GPU instance SageMaker offers that fits a 7B LoRA fine-tune in fp16/bf16 with a 4-bit
base:
# train_lora.py driver — run from a notebook or a local Python shell with the sagemaker SDK # pip install "sagemaker>=2.220,<3" (pinned; the SDK's estimator API has changed across majors) import sagemaker from sagemaker.pytorch import PyTorch role = "arn:aws:iam::<account-id>:role/<your-sagemaker-execution-role>" session = sagemaker.Session() estimator = PyTorch( entry_point="train.py", # f4-lora's own PEFT LoRA training script, unchanged source_dir="./f4-lora-src", # train.py + requirements.txt (peft, transformers, bitsandbytes, pinned) role=role, framework_version="2.3", # pinned PyTorch container version py_version="py311", instance_type="ml.g4dn.xlarge", # 1x T4 16GB — the smallest GPU instance that fits this job instance_count=1, hyperparameters={ "r": 16, "alpha": 32, "target_modules": "q_proj,v_proj", "epochs": 3, }, output_path="s3://<your-bucket>/f4-lora/artifacts", )
2. The job. fit() blocks until the job reaches a terminal state; this is
the only step that bills by the second, and only while it runs:
estimator.fit({"training": "s3://<your-bucket>/f4-lora/data/train"})
# ... CloudWatch logs stream here ...
# 2024-xx-xx INFO training job completed
# Training seconds: 5,412 Billable seconds: 5,412 -> ~1.5 h x ~$0.74/h (approx.)
3. The endpoint. This is the resource that behaves differently from the job above — it does not exit when it is done, because a real-time endpoint is never "done":
predictor = estimator.deploy(
initial_instance_count=1,
instance_type="ml.g4dn.xlarge",
endpoint_name="f4-lora-endpoint",
)
# -------------! <- SageMaker's own progress dots; this line alone means "billing has started"
4. One invocation. Prove it answers, once — this project needs exactly one real request, not a load test:
result = predictor.predict({"inputs": "Summarize: " + open("sample.txt").read()})
print(result)
# {"generated_text": "..."} -- one request; the endpoint keeps billing after this returns
5. delete_endpoint — the same session, before you do anything else.
predictor.delete_endpoint()
# or, from the CLI if the notebook session is gone:
aws sagemaker delete-endpoint --endpoint-name f4-lora-endpoint
aws sagemaker delete-endpoint-config --endpoint-config-name f4-lora-endpoint
aws sagemaker delete-model --model-name <model name from `aws sagemaker list-models`>
aws sagemaker describe-training-job --training-job-name <name> --query 'TrainingJobStatus'
prints "Completed", and BillableTimeInSeconds in the same output
matches roughly what you expected the run to take.aws s3 ls s3://<your-bucket>/f4-lora/artifacts/ lists a
model.tar.gz under the training job's name.result is generated text
from your adapter, not an error — and aws sagemaker describe-endpoint --endpoint-name
f4-lora-endpoint --query 'EndpointStatus' read "InService" before you
called it.aws sagemaker delete-endpoint --endpoint-name f4-lora-endpoint 2>/dev/null
aws sagemaker delete-endpoint-config --endpoint-config-name f4-lora-endpoint 2>/dev/null
aws sagemaker list-models --query 'Models[].ModelName' --output text | tr '\t' '\n' | grep f4-lora \
| xargs -I{} aws sagemaker delete-model --model-name {}
# stop any notebook instance you used to drive this, if one was left running
aws sagemaker list-notebook-instances --query 'NotebookInstances[?NotebookInstanceStatus==`InService`].NotebookInstanceName' --output text \
| tr '\t' '\n' | xargs -I{} aws sagemaker stop-notebook-instance --notebook-instance-name {}
aws sagemaker list-endpoints --query 'Endpoints[].EndpointName'
# []
What remains and what it costs, approx., verify on the pricing page
today: the model.tar.gz artifact and any training data in S3 (a few cents per
GB-month, the same "keep 5" instinct as cloud-05's ECR lifecycle policy), and nothing else —
the training job cannot be "deleted" because it already stopped billing the moment it
completed; only the endpoint, its config and the model registration needed an explicit delete,
and all three are gone above.
This is self-attestation — the site cannot see your AWS account, so checking the box and pressing the button is you telling The Path the endpoint answered one real request and is now gone.
Completed is also "billing over". The
endpoint's clock does not know the difference between "answering real traffic" and "sitting
there since last night" — only delete_endpoint stops it, which is exactly why this
page's mark-done will not file until you have run it and proven list-endpoints
empty.