Skip to content

Edge inference

S3 laketelemetry/
→
TrainSageMaker or local
→
ZoneAnomalyonnxruntime
→
red RGB
→
DynamoDB
Edge scores every telemetry sample. Two anomalies in a row turn the zone red — before the cloud alarm can fire.

The cloud alarm takes minutes and needs the internet. This page closes the loop with ML: the cloud learns what “normal” looks like for each zone, and the core
Greengrass core device — the Linux host running Nucleus (this lab: Intel NUC6CAY). Client devices discover and connect to it.
applies that knowledge locally within seconds. The three pieces:

PieceFileWhat it does
Training scriptartifacts/ml/zone-anomaly/train.py

Fits a one-component Gaussian mixture per zone on tempC, skipping the first 5 minutes after boot. Threshold at 4 σ; ONNX outputs score_samples (log-likelihood).

Training jobartifacts/ml/zone-anomaly/training-job.json

SageMaker scikit-learn container in script mode; S3 telemetry/ as the train channel

Edge component

ZoneAnomaly-1.0.0 + infer.py

Scores every telemetry message. Two consecutive anomalies turn the RGB red; two normal readings clear it.

SageMaker Edge Manager is not used; it was discontinued (Edge Manager EOL). A custom component plus ONNX is the path AWS describes for ML inference on Greengrass.

Terminal window
set -a && source config/walkthrough.env && set +a
export GG_EDGE_ALLOW_AWS=1
AWS_ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
ARTIFACT_BUCKET="${PROJECT_NAME}-${ENVIRONMENT}-gg-artifacts-${AWS_ACCOUNT_ID}"
DATA_BUCKET="${PROJECT_NAME}-${ENVIRONMENT}-gg-data-${AWS_ACCOUNT_ID}"
RULE_PREFIX=$(echo "${PROJECT_NAME}_${ENVIRONMENT}" | tr '-' '_')
IOT_RULE_ROLE_ARN=$(aws iam get-role --role-name "${PROJECT_NAME}-${ENVIRONMENT}-iot-rules" \
--query Role.Arn --output text)
ZONE_TABLE="${ZONE_TABLE:-${PROJECT_NAME}-${ENVIRONMENT}-zone-state}"
SAGEMAKER_ROLE="${PROJECT_NAME}-${ENVIRONMENT}-sagemaker-training"
CORE_THING_GROUP_ARN=$(aws iot describe-thing-group \
--thing-group-name "$CORE_THING_GROUP" --query thingGroupArn --output text)
(no output)

Leave the zones alone, with no thumbs or hair dryer, for at least 25 minutes after Archive telemetry started. The script needs 120 readings per zone taken after the first 5 minutes of uptime.

Terminal window
for t in $CLIENT_THING_NAME ${CLIENT_THING_NAME_2:-}; do
echo "$t $(aws s3 ls "s3://${DATA_BUCKET}/telemetry/thing=${t}/" --recursive | wc -l)"
done
gg-edge-wt-dev-esp32-1 227
gg-edge-wt-dev-esp32-2 229

If you heated a zone on the alarm page, those readings are in the lake too. A few minutes of heat among hundreds of normal samples only widens σ slightly.

The trust and permissions follow SageMaker roles but are scoped to this bucket and the scikit-learn image repository.

Pick the scikit-learn image for your Region

Section titled “Pick the scikit-learn image for your Region”

The account that hosts the prebuilt image differs per Region. Look yours up in the SageMaker Python SDK table image_uri_config/sklearn.json. For ap-southeast-2 it is 783357654285:

Terminal window
SKLEARN_REGISTRY=783357654285
SKLEARN_IMAGE_URI="${SKLEARN_REGISTRY}.dkr.ecr.${AWS_REGION}.amazonaws.com/sagemaker-scikit-learn:1.4-2-cpu-py3"
echo "$SKLEARN_IMAGE_URI"
783357654285.dkr.ecr.ap-southeast-2.amazonaws.com/sagemaker-scikit-learn:1.4-2-cpu-py3
Terminal window
aws iam create-role --role-name "$SAGEMAKER_ROLE" \
--assume-role-policy-document file://artifacts/policies/sagemaker-trust.json \
--query Role.Arn --output text
sed -e "s|DATA_BUCKET|${DATA_BUCKET}|g" -e "s|AWS_REGION|${AWS_REGION}|g" \
-e "s|SKLEARN_REGISTRY|${SKLEARN_REGISTRY}|g" \
artifacts/policies/sagemaker-training.json > /tmp/sagemaker-training.json
aws iam put-role-policy --role-name "$SAGEMAKER_ROLE" \
--policy-name zone-anomaly-training --policy-document file:///tmp/sagemaker-training.json
arn:aws:iam::123456789012:role/gg-edge-wt-dev-sagemaker-training
Terminal window
tar czf /tmp/sourcedir.tar.gz -C artifacts/ml/zone-anomaly train.py requirements.txt
aws s3 cp /tmp/sourcedir.tar.gz "s3://${DATA_BUCKET}/ml/source/sourcedir.tar.gz"
upload: ../../tmp/sourcedir.tar.gz to s3://gg-edge-wt-dev-gg-data-123456789012/ml/source/sourcedir.tar.gz

New accounts often have 0 ml.m5.large training quota. If create-training-job fails with ResourceLimitExceeded, request an increase (ml.m5.large for training job usage) in Service Quotas, or use the local fallback below — same train.py, same model.tar.gz layout.

Terminal window
TRAINING_JOB_NAME="zone-anomaly-$(date -u +%Y%m%d%H%M%S)"
sed -e "s|TRAINING_JOB_NAME|${TRAINING_JOB_NAME}|g" \
-e "s|SAGEMAKER_ROLE_ARN|arn:aws:iam::${AWS_ACCOUNT_ID}:role/${SAGEMAKER_ROLE}|g" \
-e "s|SKLEARN_IMAGE_URI|${SKLEARN_IMAGE_URI}|g" \
-e "s|DATA_BUCKET|${DATA_BUCKET}|g" -e "s|AWS_REGION|${AWS_REGION}|g" \
artifacts/ml/zone-anomaly/training-job.json > /tmp/training-job.json
aws sagemaker create-training-job --cli-input-json file:///tmp/training-job.json
{
"TrainingJobArn": "arn:aws:sagemaker:ap-southeast-2:123456789012:training-job/zone-anomaly-…"
}
Terminal window
aws sagemaker wait training-job-completed-or-stopped --training-job-name "$TRAINING_JOB_NAME"
aws sagemaker describe-training-job --training-job-name "$TRAINING_JOB_NAME" \
--query "[TrainingJobStatus,FailureReason,ModelArtifacts.S3ModelArtifacts]" --output text
Completed None s3://gg-edge-wt-dev-gg-data-123456789012/ml/output/zone-anomaly-…/output/model.tar.gz

Use Python 3.12 (3.14 cannot install the scikit-learn pin). Sync the lake, train, upload model.tar.gz where packaging expects it:

Terminal window
TRAINING_JOB_NAME="zone-anomaly-local-$(date -u +%Y%m%d%H%M%S)"
rm -rf /tmp/gg-edge-train /tmp/gg-edge-model
mkdir -p /tmp/gg-edge-train /tmp/gg-edge-model
aws s3 sync "s3://${DATA_BUCKET}/telemetry/" /tmp/gg-edge-train/telemetry/
python3.12 -m venv /tmp/gg-edge-ml-venv
/tmp/gg-edge-ml-venv/bin/pip install -U pip
/tmp/gg-edge-ml-venv/bin/pip install 'scikit-learn==1.4.2' 'numpy==1.26.4' \
-r artifacts/ml/zone-anomaly/requirements.txt
export TRAINING_JOB_NAME
export SM_CHANNEL_TRAIN=/tmp/gg-edge-train/telemetry
export SM_MODEL_DIR=/tmp/gg-edge-model
/tmp/gg-edge-ml-venv/bin/python artifacts/ml/zone-anomaly/train.py
tar czf /tmp/model.tar.gz -C /tmp/gg-edge-model .
aws s3 cp /tmp/model.tar.gz \
"s3://${DATA_BUCKET}/ml/output/${TRAINING_JOB_NAME}/output/model.tar.gz"
loaded 456 samples from 2 zones; skipped 0
trained gg-edge-wt-dev-esp32-1: samples=227 mean=30.00 sigma=0.43 alarm beyond ±1.73 C threshold=-8.082 flagged=0
trained gg-edge-wt-dev-esp32-2: samples=229 mean=27.15 sigma=0.58 alarm beyond ±2.32 C threshold=-8.374 flagged=0

5. Package the model as a Greengrass artifact

Section titled “5. Package the model as a Greengrass artifact”

The recipe unarchives zone-anomaly-model.zip into {artifacts:decompressedPath}/zone-anomaly-model/ (recipe reference).

Terminal window
MODEL_TGZ="s3://${DATA_BUCKET}/ml/output/${TRAINING_JOB_NAME}/output/model.tar.gz"
rm -rf /tmp/zone-anomaly-model /tmp/zone-anomaly-model.zip
mkdir -p /tmp/zone-anomaly-model
aws s3 cp "$MODEL_TGZ" /tmp/model.tar.gz --quiet
tar xzf /tmp/model.tar.gz -C /tmp/zone-anomaly-model
(cd /tmp/zone-anomaly-model && python3 -m zipfile -c /tmp/zone-anomaly-model.zip manifest.json *.onnx)
python3 -m zipfile -l /tmp/zone-anomaly-model.zip
File Name Modified Size
manifest.json 2026-10-10 12:14:04 837
gg-edge-wt-dev-esp32-1.onnx 2026-10-10 12:14:04 976
gg-edge-wt-dev-esp32-2.onnx 2026-10-10 12:14:04 976
Terminal window
P="s3://${ARTIFACT_BUCKET}/artifacts/com.example.ZoneAnomaly/1.0.0"
aws s3 cp artifacts/components/artifacts/com.example.ZoneAnomaly/1.0.0/infer.py "$P/infer.py"
aws s3 cp /tmp/zone-anomaly-model.zip "$P/zone-anomaly-model.zip"
upload: …/infer.py to s3://gg-edge-wt-dev-gg-artifacts-123456789012/artifacts/com.example.ZoneAnomaly/1.0.0/infer.py
upload: …/zone-anomaly-model.zip to s3://gg-edge-wt-dev-gg-artifacts-123456789012/artifacts/com.example.ZoneAnomaly/1.0.0/zone-anomaly-model.zip
Terminal window
sed "s|ARTIFACT_BUCKET|${ARTIFACT_BUCKET}|g" \
artifacts/components/recipes/com.example.ZoneAnomaly-1.0.0.yaml > /tmp/ZoneAnomaly-1.0.0.yaml
aws greengrassv2 create-component-version --inline-recipe fileb:///tmp/ZoneAnomaly-1.0.0.yaml \
--query "[componentName,componentVersion,status.componentState]" --output text
com.example.ZoneAnomaly 1.0.0 REQUESTED

Wait until describe-component shows DEPLOYABLE. After you retrain, upload under a new version path (for example 1.0.1) and bump ComponentVersion — overwriting the ZIP in place fails the integrity check.

The install step builds a virtualenv with onnxruntime. Ubuntu 24.04 needs:

Terminal window
sudo apt-get install -y python3-venv
Setting up python3.12-venv (3.12.3-1ubuntu0.17) ...

Add ZoneAnomaly to the greenhouse deployment

Section titled “Add ZoneAnomaly to the greenhouse deployment”

A Thing-group deployment replaces the previous one, so render the full stack from Deploy from the cloud again and add the new component:

Terminal window
CLI_VER=$(aws greengrassv2 list-components --scope PUBLIC \
--query "components[?componentName=='aws.greengrass.Cli'].latestVersion.componentVersion" --output text)
LM_VER=$(aws greengrassv2 list-components --scope PUBLIC \
--query "components[?componentName=='aws.greengrass.LogManager'].latestVersion.componentVersion" --output text)
[ -n "$LM_VER" ] || LM_VER=2.3.14
sed -e "s|TARGET_THING_GROUP_ARN|${CORE_THING_GROUP_ARN}|g" \
-e "s|CLIENT_THING_PREFIX|${CLIENT_THING_PREFIX}|g" \
-e "s|CLI_COMPONENT_VERSION|${CLI_VER}|g" \
-e "s|LOGMANAGER_COMPONENT_VERSION|${LM_VER}|g" \
artifacts/deployments/greenhouse-cloud.json \
| jq '.components["com.example.ZoneAnomaly"] = {"componentVersion": "1.0.0"}' \
> /tmp/greenhouse-cloud-ml.json
aws greengrassv2 create-deployment --cli-input-json file:///tmp/greenhouse-cloud-ml.json
{
"deploymentId": "…"
}

Wait for SUCCEEDED. First install downloads onnxruntime wheels (~30 s on the NUC6CAY). The recipe allows 900 s for slower links.

Terminal window
sed -e "s|ZONE_TABLE|${ZONE_TABLE}|g" -e "s|IOT_RULE_ROLE_ARN|${IOT_RULE_ROLE_ARN}|g" \
artifacts/iot-rules/inference-to-dynamodb.json > /tmp/rule-inference-ddb.json
aws iot create-topic-rule --rule-name "${RULE_PREFIX}_inference_ddb" \
--topic-rule-payload file:///tmp/rule-inference-ddb.json
(no output)
Terminal window
aws dynamodb get-item --table-name "$ZONE_TABLE" \
--key "{\"thing\":{\"S\":\"${CLIENT_THING_NAME}\"},\"kind\":{\"S\":\"inference\"}}" \
--query "Item.{tempC:tempC.N,score:score.N,anomaly:anomaly.BOOL,model:model.S}"
{
"tempC": "30.9",
"score": "-2.2566",
"anomaly": false,
"model": "zone-anomaly-local-20261009231335"
}

model is the training job name (SageMaker or zone-anomaly-local-…).

On the NUC, /greengrass/v2/logs/com.example.ZoneAnomaly.log shows loaded ONNX models and per-sample scores.

Hold your thumb on the esp32-1 module and watch the RGB:

TimeRGBWho acted
~20 sred

ZoneAnomaly on the NUC (two anomalous samples in a row)

1–3 minblueCloudWatch alarm → SNS → Lambda → IoT Core → bridge → GgEdgeLoop
after you let gooff

Edge clears after two normal samples; alarm OK sends off

After two hot readings, DynamoDB shows:

{
"tempC": "36.5",
"score": "-112.8415",
"anomaly": true,
"model": "zone-anomaly-local-20261009231335"
}
{
"state": "on",
"color": "red",
"source": "com.example.ZoneAnomaly",
"reason": "anomaly score -112.8415 at tempC 36.5"
}

The model is two-sided, so a zone that suddenly runs cold also goes red. The cloud alarm only watches the high side.

When finished: Teardown.