> For the complete documentation index, see [llms.txt](https://renewables.docs.helinplatform.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://renewables.docs.helinplatform.com/sgm-product-docs/guides/troubleshooting.md).

# Troubleshooting

Common issues hitting the SGM API and how to diagnose them, auth failures, curtailments not landing, schedules not applying, timeouts, error response codes.

A short field guide to the issues that come up most often when integrating against the SGM API. Each section covers what symptom you're seeing, what's likely wrong, and what to check first.

## 1. Authentication failures

**Symptom:** the API returns `401 Unauthorized` or `403 Forbidden`.

**What to check, in order:**

* Both the `Authorization: Bearer …` header **and** the `X-Management-Token` header are present on every request. Missing either one will fail. See [Authentication](https://app.gitbook.com/s/x327p9qmgvN6jAxpSI6V/authentication).
* The Auth0 access token hasn't expired. Tokens last 24 hours, if your integration has been running longer than that without refreshing, you'll get `401`s.
* The `client_id` and `client_secret` you used to mint the Auth0 token are still valid. Helin can reissue if either was rotated.
* For site-scoped operations, your `X-Management-Token` includes the target `site_id`. Listing tokens via `GET /v1/auth-tokens` shows which sites each token can reach.
* You're hitting the right base URL: `https://smartsolar.polarisedge.com/v1`. A `401` from a wrong host is generic and easy to misread.

A `401` typically means the token is invalid or missing. A `403` typically means the token is valid but doesn't have access to that specific resource.

## 2. Curtailment commands not working

**Symptom:** `PUT /sites/{site_id}/curtail-category` returns a successful response, but the device never actually changes value.

**What to check:**

* The device's `include_in_curtailment` flag is `true`. Inspect via `GET /v1/devices/{device_id}`. A device with `include_in_curtailment: false` is excluded from category-level curtailment by design.
* The device is in the right `category`. If you target `pv_inverter` but the device is categorised as `inverter`, it won't be hit. List the site's categories via `GET /v1/sites/{site_id}/device-categories`.
* The value is within `min_value` and `max_value`. Out-of-range values are rejected per device, not at the request level, check the `not_updated_devices` map in the response.
* The node is online. Use `GET /v1/site-health-report` and look for `healthy: false` entries.
* The device isn't backing off. After three failed writes, SGM puts the device in an exponential backoff list (skips 1, 2, 4, … up to 16 cycles). Check `GET /v1/device-error-logs?period=1` for recent errors.
* No higher-priority strategy is locking the device. Site Control strategies (Power Assist, Protective Mode) take ownership of the BESS and reject direct setpoints, see [Site Control](/sgm-product-docs/reference-docs/site-control.md).

## 3. Schedule not being applied

**Symptom:** you submitted a schedule for today, but the device isn't following it.

**What to check:**

* The `valid_on` date is correct, in `YYYY-MM-DD` format, and matches the timezone you expect (schedules are interpreted in the **site's local timezone**).
* The `category` matches a real category on the site (`GET /v1/sites/{site_id}/device-categories`).
* Times are aligned to **15-minute UTC increments** (`00:00`, `00:15`, `00:30`, `00:45`). Sub-15-minute values get rounded down.
* The schedule was posted **at least 5 minutes before** the first slot. Late submissions risk the first slot starting against the device's `default_value`. See [Scheduling setpoints](/sgm-product-docs/reference-docs/scheduling-setpoints.md#scheduled-setpoints).
* Inspect `synchronization_result` in the schedule response. Devices showing `pending` haven't acknowledged yet, wait a few minutes and re-check via `GET /v1/sites/{site_id}/schedules/{valid_on}`.
* No direct API request is overriding it. Direct setpoints win until the watchdog fires; the schedule resumes after that.

## 4. Request timeouts

**Symptom:** the API call takes a long time and eventually times out at your client.

**What to check:**

* For curtailment writes, edge devices have to acknowledge over Modbus or MQTT. A site with 50 inverters and a flaky Modbus segment will take longer than one with two. Use site-scoped requests where possible.
* Increase your client's HTTP timeout for `PUT /sites/{site_id}/curtail-category`. 30 seconds is a reasonable floor; some sites legitimately take 10-20 seconds when the bus is busy.
* Check your network path to `smartsolar.polarisedge.com`. Cloud-side latency is usually under 100 ms; if your traceroute shows 500+ ms hops, something between you and the API is slow.
* Don't retry timeouts blindly. Use idempotency-key patterns or check via `GET /v1/sites/{site_id}/logs` whether the previous attempt actually landed before resending.

## 5. Understanding error responses

| Status | What it means                                                                    | First diagnostic step                              |
| ------ | -------------------------------------------------------------------------------- | -------------------------------------------------- |
| `400`  | Validation error before the request reached the platform, usually malformed body | Inspect `detail` in the `HTTPValidationError` body |
| `401`  | Missing or invalid Auth0 token, or invalid `X-Management-Token`                  | Refresh tokens, verify both headers are present    |
| `403`  | Token is valid but lacks permission for this resource                            | Check token's site list                            |
| `404`  | Resource doesn't exist or your token can't see it                                | Verify the ID and the token's scope                |
| `409`  | State conflict, typically a Site Control strategy owns the targeted asset        | Inspect active strategies on the site              |
| `422`  | Body shape invalid                                                               | Inspect `detail[]` for the offending field         |
| `424`  | Failed dependency, gateway, IoT Hub, or downstream system unavailable            | Check `GET /v1/site-health-report`                 |
| `429`  | Rate limit exceeded (10 calls/min per site for `curtail-category`)               | Honour `Retry-After` header                        |
| `500`  | Server error, transient                                                          | Retry with exponential backoff                     |
| `502`  | Gateway unreachable                                                              | Check `GET /v1/site-health-report`                 |
| `504`  | No acknowledgement from gateway in time                                          | Retry; check Modbus health on site                 |

The full validation envelope is documented at [Errors and validation](https://app.gitbook.com/s/x327p9qmgvN6jAxpSI6V/errors).

## 6. Monitoring proactively

The two health endpoints are the cheapest insurance you can buy:

* **`GET /v1/ping`**, verifies the API itself is reachable. Returns `{"ping": "pong!"}`. No auth required, safe to hammer from monitoring.
* **`GET /v1/site-health-report`**, verifies every node in your portfolio is online. A `healthy: false` entry means commands to that site won't reach the devices.
* **`GET /v1/dataflow-health-report`**, verifies node telemetry is flowing into the data store. Useful for catching cases where a node is online but its dataflow pipeline is broken.

Pull these on a 1-minute or 5-minute cadence and alert on changes; you'll usually find out about an outage before your operators do.

## When to escalate to Helin support

Open a ticket if:

* You're seeing repeated `500` errors that don't clear with backoff.
* `site-health-report` shows nodes in `healthy: false` state for more than 15 minutes.
* `device-error-logs` shows the same Modbus error pattern across multiple devices on the same node.
* A token you minted yourself disappeared or stopped working without you deleting it.

Include the `request_id` from the failing response, it's the fastest way Helin support can find the call in their logs.

## Relevant pages

<table data-view="cards"><thead><tr><th></th><th></th></tr></thead><tbody><tr><td><p><i class="fa-key">:key:</i></p><p><a href="https://app.gitbook.com/s/x327p9qmgvN6jAxpSI6V/authentication"><strong>Authentication</strong></a></p></td><td>The two-step auth flow.</td></tr><tr><td><p><i class="fa-seal-question">:seal-question:</i></p><p><a href="https://app.gitbook.com/s/x327p9qmgvN6jAxpSI6V/errors"><strong>Errors &#x26; validation</strong></a></p></td><td>Status codes and the validation envelope.</td></tr><tr><td><p><i class="fa-heart-pulse">:heart-pulse:</i></p><p><a href="/api-reference/health.md"><strong>Health &#x26; version</strong></a></p></td><td>Endpoints for proactive monitoring.</td></tr><tr><td><p><i class="fa-rectangle-history">:rectangle-history:</i></p><p><a href="/api-reference/logs.md"><strong>Logs</strong></a></p></td><td>Operation logs and device error logs.</td></tr></tbody></table>


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://renewables.docs.helinplatform.com/sgm-product-docs/guides/troubleshooting.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
