Troubleshooting

Diagnose common self-hosted failures.

Start with the service that logged the error: docker compose logs -f gateway, ... core-api, ... logs-service, ... <plugin>-api, or ... redis. Health probes are curl http://<host>:4000/health (gateway) and curl http://<host>:3300/health (Core API); both return ok when healthy.

Gateway and supergraph

  • /graphql never becomes ready (gateway logs Waiting for plugin <name> to join service discovery or WAITING FOR: <name> graphql endpoint): the gateway waits for every name in getPlugins() (core plus both enabled lists) to register and answer _service { sdl }. Compare ENABLED_PLUGINS/ENABLED_PLUGINS_ONLY_API with the plugin APIs actually deployed; remove entries whose API is not deployed, or deploy them. MAX_PLUGIN_RETRY bounds retries; once exhausted the gateway exits with code 1.
  • Supergraph composition failed, New supergraph could not be parsed, COLLISION: a subgraph's SDL is broken. Check the rover supergraph compose output on stderr in the gateway logs and each backend's POST _service { sdl } at its registered address, then fix or remove the offending subgraph. Composition runs in every gateway replica and the router reloads in place on success.
  • A plugin registers but its GraphQL calls fail: the address stored in erxes-service-<name> is not reachable from the gateway (LOAD_BALANCER_ADDRESS was inferred as http://plugin-<name>-api:<port> in production). Set LOAD_BALANCER_ADDRESS explicitly, or DEL the stale erxes-service-<name> key and restart the plugin.
  • Supergraph stays stale after a plugin restarted: registration enqueues a gateway-update-apollo-router job delayed 10 seconds (3 attempts, exponential backoff, deduplicated by the gateway:update-apollo-router:pending lock). If the lock leaked, delete it and re-register a service.
  • EADDRINUSE :::4000 or EACCES: the host port is already bound. Free it or remap - "4001:4000"; services do not retry bind failures.

Check Redis state directly: docker compose exec redis redis-cli -a <password> KEYS 'erxes-service-*' and GET erxesservice:config:<name> (see service-discovery.ts for the key format).

Auth and sessions

  • Login works locally but fails behind HTTPS: browsers drop the secure auth-token cookie (set whenever NODE_ENV is not test/development, with a 14-day maxAge) over plain HTTP. Serve the UI over HTTPS. Auth also accepts Authorization: Bearer, erxes-core-token, erxes-app-token, and x-app-api-token headers; SAME_SITE=none enables cross-site cookies when the origin differs from DOMAIN.
  • token auth failed / Invalid signature: JWT_TOKEN_SECRET differs between the gateway, core-api, and a plugin. Align the env and restart. The code default is SECRET; never deploy with it.
  • Requests lose identity through /pl:* or /graphql: the gateway repacks auth into a base64 user header (setUserHeader in userMiddleware.ts) and warns once it exceeds 32 KB; oversized headers get dropped by proxies. Set DEBUG_GATEWAY_AUTH to log user-header-set, user-token-expired, and user-token-error per request. A missing header means the request arrived unauthenticated; check token expiry and user_token_* Redis keys.
  • Sessions expire or all users are logged out: sessions are Redis-backed (user_token_<userId>_<token> keys, set with a 24-hour TTL at login). An eviction policy, flush, or Redis restart without AOF invalidates every session. Keep noeviction and persistence.

Plugins missing or broken

  • Plugin absent from navigation: query GET <core-api>/get-frontend-plugins; if the plugin is not listed with a remoteEntry.js URL, add the base name to ENABLED_PLUGINS and restart core-api. VERSION=saas additionally filters by purchased charge.
  • remoteEntry.js 404: the URL is plugins.erxes.io/<version>/<name>_ui/remoteEntry.js, where <version> comes from the plugin's RELEASE_VERSION registration. The ci-ui-* workflow must have published assets for that version; use a RELEASE_VERSION starting with 3. or leave it unset for latest.
  • A plugin route or widget fails to load in the browser: check the DevTools network tab for the remote entry fetch and loadRemote errors in the console. The runtime remote name uses underscores (pos-client → pos_client_ui) while the CDN path keeps the original name; confirm the remoteEntry.js URL resolves publicly and the plugin's config.tsx name matches.
  • /pl:<name> returns Service not found: the service name after /pl: must match the registered name, and erxes-service-<name> must hold a reachable address; the gateway proxies /pl:<serviceName> to that address and rewrites the path to /. An empty or stale value produces the 404; fix LOAD_BALANCER_ADDRESS or delete the key and restart.

Data services

  • Transaction numbers are only allowed on a replica set member or Transaction ... requires a replica set: MONGO_URL targets a standalone mongod, but some plugins wrap writes in startSession transactions. Initialize a replica set (rs.initiate()) and include replicaSet= in the URI.
  • Backend cannot connect after fixing the URI: with directConnection=true the URI's host must be reachable from the containers; do not mix 127.0.0.1 between host and container contexts.
  • WRONGPASS / NOAUTH from Redis: REDIS_PASSWORD does not match the Redis container's requirepass. Match them; the code only reads REDIS_HOST, REDIS_PORT, REDIS_PASSWORD, with no REDIS_URL.
  • Activity logs or undo history empty: the logs service may not be running, or the <db>_logs database is missing. Logs write to a separate database (erxes → erxes_logs); the logs collection TTL-deletes after LOG_RETENTION_DAYS (default 365).
  • Automations never fire: the automations service must be running and erxes-active-plugins must list it. Automations are triggered by redisPubSub events and BullMQ queues (automations-trigger, -action, -aiAgent); they are not a GraphQL subgraph.

Networking and CORS

  • Browser CORS errors: the allow-list is exact origins (DOMAIN, WIDGETS_DOMAIN, comma-separated ALLOWED_DOMAINS) plus comma-separated regexes in ALLOWED_ORIGINS. There is no automatic subdomain wildcard; add tenant hosts via ALLOWED_DOMAINS or a regex.
  • API calls go to the wrong host: REACT_APP_API_URL on the core-ui container must be the public gateway URL the browser can reach. <subdomain> inside it is replaced by the hostname's first label at runtime.
  • Wrong tenant resolved: the subdomain is the hostname's first label, read from the nginx-hostname header, then the forwarded host; a flattened hostname (no labels) resolves to the base domain.
  • Files upload but links break: uploads are configured through Settings → file upload system configs (UPLOAD_SERVICE_TYPE, AWS_*, CLOUDFLARE_*, …), not env vars. Local FILE_SYSTEM_PUBLIC mode serves read-file?key= URLs.

After any fix, re-run docker compose ps, curl -fsS http://<host>:4000/health, the __typename GraphQL probe, a fresh login, and one write per enabled plugin. If composition still fails, run rover supergraph compose manually against the same supergraph.yaml inputs the gateway logs.

Was this helpful?