Fecha agente.md secao 54-55 (infraestrutura) e 161 (WebSocket multi-tenant).
Entrega o pipeline de push em tempo real completo — o consumo visual
("Monitoramento -> Filas/Ramais") fica pra fase Frontend.
Requisito central da secao 161 ("nao transmitir tudo e filtrar so no
browser"): RealtimeGateway tem um unico ponto de emissao,
broadcastToTenant(), sempre server.to(`tenant:<id>`), nunca broadcast
global. Cada socket entra na room do proprio tenant no handshake, nunca
escolhe a room.
Autenticacao na conexao (handshake.auth.token, nao Authorization header):
valida o JWT (mesmo verifyAccessToken do JwtAuthGuard), exige tenantId no
token e a permission monitoring.view (ja existia desde RBAC, sem
consumidor ate agora) — mesmo principio de nunca confiar em tenant_id do
client, so do JWT ja emitido por /auth/select-tenant.
Origem dos eventos: canal Redis unico b2bcall:events (o mesmo desde Event
Socket). Dois produtores: b2bcall-fs-events (eventos do FreeSWITCH,
resolvendo tenantId por fan-out quando nao ha channel variable, ver
tenant-resolve.ts) e apps/api (mudancas no nosso Agent.state via
agents-me.controller, tenantId direto do JWT, sem fan-out).
Bug real achado e corrigido ao construir esta fase: nenhum evento CUSTOM do
ESL (sofia::register, sofia::gateway_state, callcenter::info) jamais
chegava em b2bcall-fs-events nesta sessao inteira. Causa: event_json(...)
mandava "CUSTOM" como ultimo token do comando `event json`, sem subclass
depois — mod_event_socket exige os subclasses logo depois do token CUSTOM
no mesmo comando pra serem entregues. Corrigido separando PLAIN_EVENTS
(viram listener .on()) de CUSTOM_SUBCLASSES (so compoem o comando de
assinatura). Resolve as lacunas ja documentadas em docs/TRUNKS.md e
docs/AGENTS.md. De quebra, corrigido um bug de nome de campo
(CC-Agent-Status, que nao existe -> CC-Agent-State) e um segundo bug real
em trunk-sync.ts (rescan nunca descarregava gateway removido -> agora roda
`killgw` antes do rescan).
Novos tipos normalizados a partir de callcenter::info, com nomes de campo
confirmados contra uma fila real: AGENT_OFFERED_CALL, AGENT_BRIDGE_FAILED,
QUEUE_MEMBER_COUNT (chamadas esperando, secao 54), QUEUE_MEMBER_LEFT (com
cause/cancelReason e timestamps — base pra Service Level/Abandon Rate
quando CDR existir).
Verificado ponta a ponta com um client socket.io real: login/pause/resume/
logout emitindo AGENT_STATE_CHANGED; chamada de teste numa fila com agente
logado emitindo QUEUE_MEMBER_COUNT/LEFT, AGENT_OFFERED_CALL,
AGENT_BRIDGE_FAILED, AGENT_STATUS_CHANGED (CC-Agent-State correto); token
ausente/invalido desconectado na hora, sem vazar nenhum evento.
Achado sistemico durante o teste (documentado, nao corrigido nesta fase):
@@unique combinado com soft delete, sem excluir deletedAt, em
Agent/Extension/Trunk/Queue/PauseReason — nao da pra reusar numero/nome/
codigo depois de apagar. Precisa de indice unico parcial em cada um, fora
do escopo desta fase.
typecheck do workspace inteiro limpo. ~144MB de memoria total nos
containers (fs-events 44MB, fs-config 45MB, freeswitch 26MB, postgres
21MB, redis 8MB).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1HxY46WGU4G1zmVDNKcWw
- dialplan_extensions table (tenant-scoped, RLS): structured editor per
agente.md secao 43 -- context, condition field/expr, actions/anti-actions
(JSON), continue, order, enabled. One condition per extension (deliberate
simplification vs raw FreeSWITCH's multi-condition extensions).
- dialplan_versions table (tenant-scoped, RLS): generate/validate/version/
activate flow (secao 44). Reactivating an older version IS the rollback
mechanism -- no separate endpoint needed.
- apps/api/src/dialplan: extensions CRUD + versions/generate (builds XML,
validates well-formedness with fast-xml-parser, saves as DRAFT) +
versions/:id/activate (atomically flips ACTIVE, supersedes the previous
one). Reused freeswitch.view/.configure permissions rather than inventing
new ones not in the agente.md permission list.
- packages/telephony: buildDialplanXml() plus ALLOWED_DIALPLAN_APPLICATIONS,
an explicit allowlist (answer/bridge/playback/hangup/set/export/... --
deliberately no system/exec/socket) guarding against a tenant configuring
a dialplan action that runs arbitrary commands on the FreeSWITCH host
(agente.md secao 180)
- b2bcall-fs-config resolves dialplan dynamically per call (unlike Trunks'
file+rescan approach -- dialplan is fetched fresh via mod_xml_curl on
every call anyway) by tenant id from the variable_b2bcall_tenant_id
channel variable already injected at directory resolution, then serving
whichever DialplanVersion is ACTIVE for that context
- verified end-to-end: created a rule for destination_number 7000, generated
and activated v1, originated a call that actually routed through the
dialplan (not bypassing it via &app()) -- CALL_CREATED -> CALL_ANSWERED ->
CALL_ENDED with the correct tenantId throughout. Created and activated a
v2, then rolled back to v1 by reactivating it; status transitions
(ACTIVE/SUPERSEDED) all confirmed via the API.
CRITICAL FINDING, fixed in this same phase: deliberately testing that the
application allowlist rejects 'system' got back 201 instead of 400 --
NestJS's ValidationPipe had been silently inert across all of apps/api's
@Body() DTOs since the API was first created. Root cause: running via
(esbuild) instead of a real build -- esbuild doesn't always
resolve cross-file parameter types for design:paramtypes metadata, and Nest
skips validation without any error when it can't determine the DTO class.
Fixed by always building with tsc before running (tsc && tsx dist/main.js
-- still via tsx because internal workspace packages aren't built to JS
yet). Re-verified with two deliberate bad-input tests post-fix, both
correctly rejected with 400. A stray malicious test row (dialplan action
'system') created while the bug was live was deleted; it was never baked
into an activated version, so nothing could have executed it.
See docs/VALIDATION_PIPE_BUG.md for the full writeup.
docs/DIALPLAN.md, docs/VALIDATION_PIPE_BUG.md, docs/EXTENSIONS.md updated
- POST /auth/login, /auth/refresh, /auth/logout, /auth/select-tenant,
/auth/change-password, GET /auth/tenants — wired to packages/auth
- JwtAuthGuard + DomainExceptionFilter (401/403 without leaking internals)
- LoginRateLimitGuard: Redis-backed 5/min per IP and per email (agente.md
secao 149), safe across multiple API instances
- helmet + restrictive cors (deny-by-default) + global rate limit
- GET /health, /health/live, /health/ready checking Postgres and Redis
- changePassword() added to packages/auth for the mustChangePassword flow
- fixed REDIS_HOST/POSTGRES_HOST docker-compose-only hostnames not
resolving from the host process; added REDIS_URL for host-side use
- verified end-to-end with curl: login, wrong password / unknown email
(same generic error), authenticated route, missing token, refresh
rotation, logout revocation, and the 429 rate limit kicking in after 5
attempts