Edge inference
The cloud alarm takes minutes and needs the internet. This page closes the
loop with ML: the cloud learns what “normal” looks like for each zone, and the
core
Greengrass core device — the Linux host running Nucleus (this lab: Intel NUC6CAY). Client devices discover and connect to it. applies that knowledge locally
within seconds. The three pieces:
| Piece | File | What it does |
|---|---|---|
| Training script | artifacts/ml/zone-anomaly/train.py | Fits a one-component Gaussian mixture per zone on |
| Training job | artifacts/ml/zone-anomaly/training-job.json | SageMaker scikit-learn container in script mode; S3
|
| Edge component |
| Scores every telemetry message. Two consecutive anomalies turn the RGB red; two normal readings clear it. |
SageMaker Edge Manager is not used; it was discontinued (Edge Manager EOL). A custom component plus ONNX is the path AWS describes for ML inference on Greengrass.
1. Session variables
Section titled “1. Session variables”set -a && source config/walkthrough.env && set +aexport GG_EDGE_ALLOW_AWS=1AWS_ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)ARTIFACT_BUCKET="${PROJECT_NAME}-${ENVIRONMENT}-gg-artifacts-${AWS_ACCOUNT_ID}"DATA_BUCKET="${PROJECT_NAME}-${ENVIRONMENT}-gg-data-${AWS_ACCOUNT_ID}"RULE_PREFIX=$(echo "${PROJECT_NAME}_${ENVIRONMENT}" | tr '-' '_')IOT_RULE_ROLE_ARN=$(aws iam get-role --role-name "${PROJECT_NAME}-${ENVIRONMENT}-iot-rules" \ --query Role.Arn --output text)ZONE_TABLE="${ZONE_TABLE:-${PROJECT_NAME}-${ENVIRONMENT}-zone-state}"SAGEMAKER_ROLE="${PROJECT_NAME}-${ENVIRONMENT}-sagemaker-training"CORE_THING_GROUP_ARN=$(aws iot describe-thing-group \ --thing-group-name "$CORE_THING_GROUP" --query thingGroupArn --output text)(no output)2. Enough normal data
Section titled “2. Enough normal data”Leave the zones alone, with no thumbs or hair dryer, for at least 25 minutes after Archive telemetry started. The script needs 120 readings per zone taken after the first 5 minutes of uptime.
for t in $CLIENT_THING_NAME ${CLIENT_THING_NAME_2:-}; do echo "$t $(aws s3 ls "s3://${DATA_BUCKET}/telemetry/thing=${t}/" --recursive | wc -l)"donegg-edge-wt-dev-esp32-1 227gg-edge-wt-dev-esp32-2 229If you heated a zone on the alarm page, those readings are in the lake too. A few minutes of heat among hundreds of normal samples only widens σ slightly.
3. SageMaker execution role
Section titled “3. SageMaker execution role”The trust and permissions follow SageMaker roles but are scoped to this bucket and the scikit-learn image repository.
Pick the scikit-learn image for your Region
Section titled “Pick the scikit-learn image for your Region”The account that hosts the prebuilt image differs per Region. Look yours up in
the SageMaker Python SDK table
image_uri_config/sklearn.json.
For ap-southeast-2 it is 783357654285:
SKLEARN_REGISTRY=783357654285SKLEARN_IMAGE_URI="${SKLEARN_REGISTRY}.dkr.ecr.${AWS_REGION}.amazonaws.com/sagemaker-scikit-learn:1.4-2-cpu-py3"echo "$SKLEARN_IMAGE_URI"783357654285.dkr.ecr.ap-southeast-2.amazonaws.com/sagemaker-scikit-learn:1.4-2-cpu-py3Create the role
Section titled “Create the role”aws iam create-role --role-name "$SAGEMAKER_ROLE" \ --assume-role-policy-document file://artifacts/policies/sagemaker-trust.json \ --query Role.Arn --output textsed -e "s|DATA_BUCKET|${DATA_BUCKET}|g" -e "s|AWS_REGION|${AWS_REGION}|g" \ -e "s|SKLEARN_REGISTRY|${SKLEARN_REGISTRY}|g" \ artifacts/policies/sagemaker-training.json > /tmp/sagemaker-training.jsonaws iam put-role-policy --role-name "$SAGEMAKER_ROLE" \ --policy-name zone-anomaly-training --policy-document file:///tmp/sagemaker-training.jsonarn:aws:iam::123456789012:role/gg-edge-wt-dev-sagemaker-training4. Train
Section titled “4. Train”Upload the training source
Section titled “Upload the training source”tar czf /tmp/sourcedir.tar.gz -C artifacts/ml/zone-anomaly train.py requirements.txtaws s3 cp /tmp/sourcedir.tar.gz "s3://${DATA_BUCKET}/ml/source/sourcedir.tar.gz"upload: ../../tmp/sourcedir.tar.gz to s3://gg-edge-wt-dev-gg-data-123456789012/ml/source/sourcedir.tar.gzStart a SageMaker job
Section titled “Start a SageMaker job”New accounts often have 0 ml.m5.large training quota. If
create-training-job fails with ResourceLimitExceeded, request an increase
(ml.m5.large for training job usage) in Service Quotas, or use the local
fallback below — same train.py, same model.tar.gz layout.
TRAINING_JOB_NAME="zone-anomaly-$(date -u +%Y%m%d%H%M%S)"sed -e "s|TRAINING_JOB_NAME|${TRAINING_JOB_NAME}|g" \ -e "s|SAGEMAKER_ROLE_ARN|arn:aws:iam::${AWS_ACCOUNT_ID}:role/${SAGEMAKER_ROLE}|g" \ -e "s|SKLEARN_IMAGE_URI|${SKLEARN_IMAGE_URI}|g" \ -e "s|DATA_BUCKET|${DATA_BUCKET}|g" -e "s|AWS_REGION|${AWS_REGION}|g" \ artifacts/ml/zone-anomaly/training-job.json > /tmp/training-job.jsonaws sagemaker create-training-job --cli-input-json file:///tmp/training-job.json{ "TrainingJobArn": "arn:aws:sagemaker:ap-southeast-2:123456789012:training-job/zone-anomaly-…"}aws sagemaker wait training-job-completed-or-stopped --training-job-name "$TRAINING_JOB_NAME"aws sagemaker describe-training-job --training-job-name "$TRAINING_JOB_NAME" \ --query "[TrainingJobStatus,FailureReason,ModelArtifacts.S3ModelArtifacts]" --output textCompleted None s3://gg-edge-wt-dev-gg-data-123456789012/ml/output/zone-anomaly-…/output/model.tar.gzLocal fallback (same script)
Section titled “Local fallback (same script)”Use Python 3.12 (3.14 cannot install the scikit-learn pin). Sync the lake,
train, upload model.tar.gz where packaging expects it:
TRAINING_JOB_NAME="zone-anomaly-local-$(date -u +%Y%m%d%H%M%S)"rm -rf /tmp/gg-edge-train /tmp/gg-edge-modelmkdir -p /tmp/gg-edge-train /tmp/gg-edge-modelaws s3 sync "s3://${DATA_BUCKET}/telemetry/" /tmp/gg-edge-train/telemetry/python3.12 -m venv /tmp/gg-edge-ml-venv/tmp/gg-edge-ml-venv/bin/pip install -U pip/tmp/gg-edge-ml-venv/bin/pip install 'scikit-learn==1.4.2' 'numpy==1.26.4' \ -r artifacts/ml/zone-anomaly/requirements.txtexport TRAINING_JOB_NAMEexport SM_CHANNEL_TRAIN=/tmp/gg-edge-train/telemetryexport SM_MODEL_DIR=/tmp/gg-edge-model/tmp/gg-edge-ml-venv/bin/python artifacts/ml/zone-anomaly/train.pytar czf /tmp/model.tar.gz -C /tmp/gg-edge-model .aws s3 cp /tmp/model.tar.gz \ "s3://${DATA_BUCKET}/ml/output/${TRAINING_JOB_NAME}/output/model.tar.gz"loaded 456 samples from 2 zones; skipped 0trained gg-edge-wt-dev-esp32-1: samples=227 mean=30.00 sigma=0.43 alarm beyond ±1.73 C threshold=-8.082 flagged=0trained gg-edge-wt-dev-esp32-2: samples=229 mean=27.15 sigma=0.58 alarm beyond ±2.32 C threshold=-8.374 flagged=05. Package the model as a Greengrass artifact
Section titled “5. Package the model as a Greengrass artifact”Repack model.tar.gz as a ZIP
Section titled “Repack model.tar.gz as a ZIP”The recipe unarchives zone-anomaly-model.zip into
{artifacts:decompressedPath}/zone-anomaly-model/
(recipe reference).
MODEL_TGZ="s3://${DATA_BUCKET}/ml/output/${TRAINING_JOB_NAME}/output/model.tar.gz"rm -rf /tmp/zone-anomaly-model /tmp/zone-anomaly-model.zipmkdir -p /tmp/zone-anomaly-modelaws s3 cp "$MODEL_TGZ" /tmp/model.tar.gz --quiettar xzf /tmp/model.tar.gz -C /tmp/zone-anomaly-model(cd /tmp/zone-anomaly-model && python3 -m zipfile -c /tmp/zone-anomaly-model.zip manifest.json *.onnx)python3 -m zipfile -l /tmp/zone-anomaly-model.zipFile Name Modified Sizemanifest.json 2026-10-10 12:14:04 837gg-edge-wt-dev-esp32-1.onnx 2026-10-10 12:14:04 976gg-edge-wt-dev-esp32-2.onnx 2026-10-10 12:14:04 976Upload model and inference code
Section titled “Upload model and inference code”P="s3://${ARTIFACT_BUCKET}/artifacts/com.example.ZoneAnomaly/1.0.0"aws s3 cp artifacts/components/artifacts/com.example.ZoneAnomaly/1.0.0/infer.py "$P/infer.py"aws s3 cp /tmp/zone-anomaly-model.zip "$P/zone-anomaly-model.zip"upload: …/infer.py to s3://gg-edge-wt-dev-gg-artifacts-123456789012/artifacts/com.example.ZoneAnomaly/1.0.0/infer.pyupload: …/zone-anomaly-model.zip to s3://gg-edge-wt-dev-gg-artifacts-123456789012/artifacts/com.example.ZoneAnomaly/1.0.0/zone-anomaly-model.zipCreate the component version
Section titled “Create the component version”sed "s|ARTIFACT_BUCKET|${ARTIFACT_BUCKET}|g" \ artifacts/components/recipes/com.example.ZoneAnomaly-1.0.0.yaml > /tmp/ZoneAnomaly-1.0.0.yamlaws greengrassv2 create-component-version --inline-recipe fileb:///tmp/ZoneAnomaly-1.0.0.yaml \ --query "[componentName,componentVersion,status.componentState]" --output textcom.example.ZoneAnomaly 1.0.0 REQUESTEDWait until describe-component shows DEPLOYABLE. After you retrain, upload
under a new version path (for example 1.0.1) and bump ComponentVersion —
overwriting the ZIP in place fails the integrity check.
6. Deploy to the core
Section titled “6. Deploy to the core”Python venv support on the NUC
Section titled “Python venv support on the NUC”The install step builds a virtualenv with onnxruntime. Ubuntu 24.04 needs:
sudo apt-get install -y python3-venvSetting up python3.12-venv (3.12.3-1ubuntu0.17) ...Add ZoneAnomaly to the greenhouse deployment
Section titled “Add ZoneAnomaly to the greenhouse deployment”A Thing-group deployment replaces the previous one, so render the full stack from Deploy from the cloud again and add the new component:
CLI_VER=$(aws greengrassv2 list-components --scope PUBLIC \ --query "components[?componentName=='aws.greengrass.Cli'].latestVersion.componentVersion" --output text)LM_VER=$(aws greengrassv2 list-components --scope PUBLIC \ --query "components[?componentName=='aws.greengrass.LogManager'].latestVersion.componentVersion" --output text)[ -n "$LM_VER" ] || LM_VER=2.3.14sed -e "s|TARGET_THING_GROUP_ARN|${CORE_THING_GROUP_ARN}|g" \ -e "s|CLIENT_THING_PREFIX|${CLIENT_THING_PREFIX}|g" \ -e "s|CLI_COMPONENT_VERSION|${CLI_VER}|g" \ -e "s|LOGMANAGER_COMPONENT_VERSION|${LM_VER}|g" \ artifacts/deployments/greenhouse-cloud.json \ | jq '.components["com.example.ZoneAnomaly"] = {"componentVersion": "1.0.0"}' \ > /tmp/greenhouse-cloud-ml.jsonaws greengrassv2 create-deployment --cli-input-json file:///tmp/greenhouse-cloud-ml.json{ "deploymentId": "…"}Wait for SUCCEEDED. First install downloads onnxruntime wheels (~30 s on
the NUC6CAY). The recipe allows 900 s for slower links.
Record every edge score in DynamoDB
Section titled “Record every edge score in DynamoDB”sed -e "s|ZONE_TABLE|${ZONE_TABLE}|g" -e "s|IOT_RULE_ROLE_ARN|${IOT_RULE_ROLE_ARN}|g" \ artifacts/iot-rules/inference-to-dynamodb.json > /tmp/rule-inference-ddb.jsonaws iot create-topic-rule --rule-name "${RULE_PREFIX}_inference_ddb" \ --topic-rule-payload file:///tmp/rule-inference-ddb.json(no output)7. Prove it
Section titled “7. Prove it”Scores in DynamoDB
Section titled “Scores in DynamoDB”aws dynamodb get-item --table-name "$ZONE_TABLE" \ --key "{\"thing\":{\"S\":\"${CLIENT_THING_NAME}\"},\"kind\":{\"S\":\"inference\"}}" \ --query "Item.{tempC:tempC.N,score:score.N,anomaly:anomaly.BOOL,model:model.S}"{ "tempC": "30.9", "score": "-2.2566", "anomaly": false, "model": "zone-anomaly-local-20261009231335"}model is the training job name (SageMaker or zone-anomaly-local-…).
On the NUC, /greengrass/v2/logs/com.example.ZoneAnomaly.log shows loaded
ONNX models and per-sample scores.
Edge vs cloud, side by side
Section titled “Edge vs cloud, side by side”Hold your thumb on the esp32-1 module and watch the RGB:
| Time | RGB | Who acted |
|---|---|---|
| ~20 s | red |
|
| 1–3 min | blue | CloudWatch alarm → SNS → Lambda → IoT Core → bridge → GgEdgeLoop |
| after you let go | off | Edge clears after two normal samples; alarm |
After two hot readings, DynamoDB shows:
{ "tempC": "36.5", "score": "-112.8415", "anomaly": true, "model": "zone-anomaly-local-20261009231335"}{ "state": "on", "color": "red", "source": "com.example.ZoneAnomaly", "reason": "anomaly score -112.8415 at tempC 36.5"}The model is two-sided, so a zone that suddenly runs cold also goes red. The cloud alarm only watches the high side.
When finished: Teardown.