Fecha agente.md secao 54-55 (infraestrutura) e 161 (WebSocket multi-tenant).
Entrega o pipeline de push em tempo real completo — o consumo visual
("Monitoramento -> Filas/Ramais") fica pra fase Frontend.
Requisito central da secao 161 ("nao transmitir tudo e filtrar so no
browser"): RealtimeGateway tem um unico ponto de emissao,
broadcastToTenant(), sempre server.to(`tenant:<id>`), nunca broadcast
global. Cada socket entra na room do proprio tenant no handshake, nunca
escolhe a room.
Autenticacao na conexao (handshake.auth.token, nao Authorization header):
valida o JWT (mesmo verifyAccessToken do JwtAuthGuard), exige tenantId no
token e a permission monitoring.view (ja existia desde RBAC, sem
consumidor ate agora) — mesmo principio de nunca confiar em tenant_id do
client, so do JWT ja emitido por /auth/select-tenant.
Origem dos eventos: canal Redis unico b2bcall:events (o mesmo desde Event
Socket). Dois produtores: b2bcall-fs-events (eventos do FreeSWITCH,
resolvendo tenantId por fan-out quando nao ha channel variable, ver
tenant-resolve.ts) e apps/api (mudancas no nosso Agent.state via
agents-me.controller, tenantId direto do JWT, sem fan-out).
Bug real achado e corrigido ao construir esta fase: nenhum evento CUSTOM do
ESL (sofia::register, sofia::gateway_state, callcenter::info) jamais
chegava em b2bcall-fs-events nesta sessao inteira. Causa: event_json(...)
mandava "CUSTOM" como ultimo token do comando `event json`, sem subclass
depois — mod_event_socket exige os subclasses logo depois do token CUSTOM
no mesmo comando pra serem entregues. Corrigido separando PLAIN_EVENTS
(viram listener .on()) de CUSTOM_SUBCLASSES (so compoem o comando de
assinatura). Resolve as lacunas ja documentadas em docs/TRUNKS.md e
docs/AGENTS.md. De quebra, corrigido um bug de nome de campo
(CC-Agent-Status, que nao existe -> CC-Agent-State) e um segundo bug real
em trunk-sync.ts (rescan nunca descarregava gateway removido -> agora roda
`killgw` antes do rescan).
Novos tipos normalizados a partir de callcenter::info, com nomes de campo
confirmados contra uma fila real: AGENT_OFFERED_CALL, AGENT_BRIDGE_FAILED,
QUEUE_MEMBER_COUNT (chamadas esperando, secao 54), QUEUE_MEMBER_LEFT (com
cause/cancelReason e timestamps — base pra Service Level/Abandon Rate
quando CDR existir).
Verificado ponta a ponta com um client socket.io real: login/pause/resume/
logout emitindo AGENT_STATE_CHANGED; chamada de teste numa fila com agente
logado emitindo QUEUE_MEMBER_COUNT/LEFT, AGENT_OFFERED_CALL,
AGENT_BRIDGE_FAILED, AGENT_STATUS_CHANGED (CC-Agent-State correto); token
ausente/invalido desconectado na hora, sem vazar nenhum evento.
Achado sistemico durante o teste (documentado, nao corrigido nesta fase):
@@unique combinado com soft delete, sem excluir deletedAt, em
Agent/Extension/Trunk/Queue/PauseReason — nao da pra reusar numero/nome/
codigo depois de apagar. Precisa de indice unico parcial em cada um, fora
do escopo desta fase.
typecheck do workspace inteiro limpo. ~144MB de memoria total nos
containers (fs-events 44MB, fs-config 45MB, freeswitch 26MB, postgres
21MB, redis 8MB).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1HxY46WGU4G1zmVDNKcWw
Achado ao iniciar a fase de Realtime Monitoring: nenhum evento CUSTOM
(sofia::register, sofia::gateway_state, callcenter::info) jamais chegou em
b2bcall-fs-events nesta sessao, apesar do estado real do FreeSWITCH mudar de
verdade (confirmado originando uma chamada de teste pra dentro de uma fila
real com agente logado). Isso deixava dois gaps documentados como "nao
verificado" em docs/TRUNKS.md e docs/AGENTS.md.
Causa raiz: `event_json(...SUBSCRIBED_EVENTS)` mandava "CUSTOM" como ultimo
token do comando `event json`, sem nenhum subclass depois. O
mod_event_socket do FreeSWITCH exige que os subclasses (callcenter::info,
sofia::register, ...) venham imediatamente depois do token CUSTOM no mesmo
comando — sem isso, zero eventos CUSTOM sao entregues, de qualquer
subclass.
Corrigido separando PLAIN_EVENTS (nomes normais, cada um vira um listener
.on()) de CUSTOM_SUBCLASSES (so compoe o comando de subscricao — o client
ESL sempre emite "CUSTOM" como nome de evento, com o subclass real no
header Event-Subclass). Corrigido tambem um bug de nome de campo:
normalizeCustomEvent lia CC-Agent-Status, que nao existe; o campo real e
CC-Agent-State.
Com o pipeline corrigido, chegam eventos ricos de callcenter::info nunca
antes observados: agent-offering, bridge-agent-fail, members-count (fila em
tempo real) e member-queue-end (com CC-Cause/CC-Cancel-Reason e timestamps
de entrada/saida — atendida vs. abandonada). Adicionados como novos tipos
normalizados: AGENT_OFFERED_CALL, AGENT_BRIDGE_FAILED, QUEUE_MEMBER_COUNT,
QUEUE_MEMBER_LEFT.
De quebra, achado e corrigido um segundo bug real ao reverificar Trunks com
o pipeline de eventos funcionando: `sofia profile external rescan` nunca
descarregava um gateway cujo arquivo .xml foi apagado (fica fantasma na
memoria do Sofia indefinidamente). trunk-sync.ts agora roda `sofia profile
external killgw <nome>` pra cada gateway removido, antes do rescan.
Reverificado ponta a ponta pra ambos os bugs:
- Trunk com register:true apontando pra host inexistente: Trunk.status no
banco passa de UNKNOWN pra FAILED sozinho, via evento, sem polling.
- Trunk apagado via API: gateway some imediatamente de `sofia status
gateway`, sem esperar reinicio de profile.
- Chamada de teste real numa fila com agente logado: members-count,
agent-offering, agent-state-change (CC-Agent-State correto: Waiting/
Receiving), bridge-agent-fail, todos chegando certos no canal Redis
b2bcall:events.
typecheck do workspace inteiro limpo.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1HxY46WGU4G1zmVDNKcWw
Fecha agente.md secao 45-49/52. Depois desta fase, um usuario autenticado
consegue logar como agente, entrar numa fila real, se pausar e voltar,
tudo refletido de verdade no FreeSWITCH.
Schema (migration 20260828124245_agents):
- agents (tenant-scoped, RLS): User -> Extension -> identidade de agente,
state (enum AgentState de 8 valores) espelhando o estado real, so
alterado via login/logout/pause/resume, nunca escrito direto pela API.
- tiers: Queue<->Agent (level/position 1:1 com mod_callcenter).
- agent_sessions: um ciclo login->logout por linha.
- agent_state_events: historico de transicoes de estado.
- pause_reasons / agent_pause_events (secao 48).
Dois bugs reais corrigidos em FreeSwitchTelephonyProvider, presentes desde
a fase de Event Socket original:
- "queue add/del member" nao existe no mod_callcenter — membership de
fila usa tier add/tier del. So foi pego agora ao confirmar de novo a
sintaxe via `help callcenter_config` antes de codar esta fase.
- addAgent/removeAgent nao existiam ainda (agent add/del).
Mecanismo de sync: agentes e tiers nao tem representacao em XML, so
comando ESL direto — diferente do padrao "regenera todos os arquivos"
usado em Trunks/Queues. apps/api publica uma mensagem por acao com
payload (b2bcall:agents:sync, b2bcall:tiers:sync); b2bcall-fs-config
aplica o comando correspondente (agent-sync.ts).
Achados confirmados manualmente contra o FreeSWITCH real antes de codar:
- `agent add`/`tier add` nao sao idempotentes (erro em duplicata) — sync
ignora esse erro (.catch), condicao esperada em resync.
- `agent del`/`tier del` em algo inexistente nao da erro — seguro chamar
sem checar existencia antes.
- `agent set status` so aceita 3 valores exatos (Available/On Break/
Logged Out) — testado deliberadamente com valor invalido.
- Corrida real: atribuir tier antes do primeiro login do agente falha
silenciosamente do lado do FreeSWITCH (agente so existe la a partir do
`agent add` no login). Login sempre re-sincroniza todos os tiers do
agente depois de garantir que ele existe — auto-correcao confirmada no
teste ponta a ponta.
apps/api: AgentsController (CRUD), AgentsMeController (login/logout/
pause/resume — sempre resolve o agente via JWT, nunca um agentId
arbitrario do client), PauseReasonsController (CRUD), QueueAgentsController
(POST/DELETE de tier em /queues/:id/agents).
Verificado ponta a ponta via curl + fs_cli contra o FreeSWITCH real:
login -> Available, pause -> On Break, resume -> Available, logout ->
Logged Out, todos batendo entre Agent.state (banco) e `agent list`
(FreeSWITCH). typecheck do workspace inteiro limpo. ~350MB de memoria
total (docker stats).
Documentado em docs/AGENTS.md, incluindo lacuna conhecida: estados
derivados de chamada (RINGING/IN_CALL/WRAP_UP/RESERVED) dependem do
evento CUSTOM callcenter::info, ainda nao comprovado chegando em
fs-events nesta sessao (mesma lacuna de sofia::gateway_state ja
documentada em docs/TRUNKS.md) — precisa de uma chamada real passando
pela fila pra investigar.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1HxY46WGU4G1zmVDNKcWw
- Investigated the real callcenter_config command surface via
'help callcenter_config' on the running FreeSWITCH before writing any
code: queues only have load/unload/reload (static XML + reload, no
'queue add' exists), while agents and tiers are fully dynamic via ESL
commands (agent add, tier add) -- no file involved. This shapes the next
phase (Agents/Tiers) differently from this one.
- queues table (tenant-scoped, RLS): strategy, moh/announce, wait times,
tier rules, discard/abandoned handling, skip-agents-with-external-calls,
recording_enabled
- packages/telephony: buildQueueXml()
- infrastructure/freeswitch: our own callcenter.conf.xml override (empties
the vanilla static agents/tiers -- those become fully dynamic in the next
phase) that includes callcenter_queues.conf.d/*.xml via X-PRE-PROCESS,
same pattern as the Sofia gateway directory
- apps/api/src/queues: CRUD (POST/GET/GET:id/DELETE) using the queues.view/
.manage permissions already in the seed
- b2bcall-fs-config (queue-sync.ts): one XML file per queue on a shared
volume, synced via Redis pub/sub (b2bcall:queues:sync) on create/delete
and once at boot -- same shape as trunk-sync.ts
- confirmed manually against the real FreeSWITCH, before coding the sync
logic: 'queue load <name>' fails ('Invalid Queue not found!') for a file
added after boot -- needs 'reloadxml' first to repopulate the in-memory
XML tree from disk; after that, 'queue reload <name>' alone handles both
create and update, no need to distinguish load vs reload
- verified end-to-end: created a queue (ROUND_ROBIN, maxWaitTime=120,
discardAbandonedAfter=90), 'callcenter_config queue list' showed the
correct values on the FreeSWITCH side; deleted it, list went back empty
docs/QUEUES.md
- dialplan_extensions table (tenant-scoped, RLS): structured editor per
agente.md secao 43 -- context, condition field/expr, actions/anti-actions
(JSON), continue, order, enabled. One condition per extension (deliberate
simplification vs raw FreeSWITCH's multi-condition extensions).
- dialplan_versions table (tenant-scoped, RLS): generate/validate/version/
activate flow (secao 44). Reactivating an older version IS the rollback
mechanism -- no separate endpoint needed.
- apps/api/src/dialplan: extensions CRUD + versions/generate (builds XML,
validates well-formedness with fast-xml-parser, saves as DRAFT) +
versions/:id/activate (atomically flips ACTIVE, supersedes the previous
one). Reused freeswitch.view/.configure permissions rather than inventing
new ones not in the agente.md permission list.
- packages/telephony: buildDialplanXml() plus ALLOWED_DIALPLAN_APPLICATIONS,
an explicit allowlist (answer/bridge/playback/hangup/set/export/... --
deliberately no system/exec/socket) guarding against a tenant configuring
a dialplan action that runs arbitrary commands on the FreeSWITCH host
(agente.md secao 180)
- b2bcall-fs-config resolves dialplan dynamically per call (unlike Trunks'
file+rescan approach -- dialplan is fetched fresh via mod_xml_curl on
every call anyway) by tenant id from the variable_b2bcall_tenant_id
channel variable already injected at directory resolution, then serving
whichever DialplanVersion is ACTIVE for that context
- verified end-to-end: created a rule for destination_number 7000, generated
and activated v1, originated a call that actually routed through the
dialplan (not bypassing it via &app()) -- CALL_CREATED -> CALL_ANSWERED ->
CALL_ENDED with the correct tenantId throughout. Created and activated a
v2, then rolled back to v1 by reactivating it; status transitions
(ACTIVE/SUPERSEDED) all confirmed via the API.
CRITICAL FINDING, fixed in this same phase: deliberately testing that the
application allowlist rejects 'system' got back 201 instead of 400 --
NestJS's ValidationPipe had been silently inert across all of apps/api's
@Body() DTOs since the API was first created. Root cause: running via
(esbuild) instead of a real build -- esbuild doesn't always
resolve cross-file parameter types for design:paramtypes metadata, and Nest
skips validation without any error when it can't determine the DTO class.
Fixed by always building with tsc before running (tsc && tsx dist/main.js
-- still via tsx because internal workspace packages aren't built to JS
yet). Re-verified with two deliberate bad-input tests post-fix, both
correctly rejected with 400. A stray malicious test row (dialplan action
'system') created while the bug was live was deleted; it was never baked
into an activated version, so nothing could have executed it.
See docs/VALIDATION_PIPE_BUG.md for the full writeup.
docs/DIALPLAN.md, docs/VALIDATION_PIPE_BUG.md, docs/EXTENSIONS.md updated
- trunks table (tenant-scoped, RLS): host/proxy/realm, register,
username/password_enc (AES-256-GCM), dtmf_mode, ping, transport, and a
status/status_updated_at pair meant to be driven by FreeSWITCH events
- apps/api/src/trunks: CRUD (POST/GET/GET:id/DELETE), same RBAC/tenant
pattern as Extensions, password never exposed in any GET
- packages/telephony: buildGatewayXml() generates a Sofia gateway XML file
- b2bcall-fs-config now writes sip_profiles/external/<trunk_id>.xml (shared
Docker volume with FreeSWITCH -- the vanilla external profile already
includes external/*.xml) and runs 'sofia profile external rescan' over
ESL; syncs on boot and on demand via Redis pub/sub
(b2bcall:trunks:sync), since apps/api runs on the host and fs-config has
no port published to reach directly
- added FreeSwitchTelephonyProvider.waitUntilConnected() to fix a startup
race: the first sync ran before the ESL connection had settled, logging
a harmless but noisy error
- verified end-to-end with a fake host: create trunk -> gateway file
written -> FreeSWITCH shows the real gateway (FAIL_WAIT, expected) ->
delete -> file removed (cleanup also correctly swept the stale
'example.com' gateway that had been copied into the volume from the
vanilla image)
- apps/freeswitch-events/src/trunk-status.ts: written to update Trunk.status
from sofia::gateway_state events, using the same normalizeEslEvent path
already proven for CHANNEL_* events
- KNOWN GAP, documented rather than glossed over: monitored fs-events for
~90s while the gateway visibly transitioned states in FreeSWITCH
(FAIL_WAIT/DOWN) and no sofia::gateway_state event was observed arriving.
CUSTOM/sofia::* events have not actually been proven working end-to-end
in this session -- only CHANNEL_* events have been. Needs verification
against a real SIP target before the status auto-update can be trusted
in production. See docs/TRUNKS.md and TODO.md.
- docs/TRUNKS.md
- extensions table (tenant-scoped, RLS): number, sip_password_enc
(AES-256-GCM via packages/shared/src/crypto.ts), caller_id, context,
sofia_profile, codecs, max_registrations
- apps/api/src/extensions: CRUD (POST/GET/GET:id/DELETE), protected by a
new generic PermissionGuard (@RequirePermission decorator), tenant
resolved only from the JWT (never trusted from the client)
- SIP password is returned in plaintext only once, in the create response;
toPublicExtension() explicitly destructures the encrypted field out
(not a spread) so it can't leak by accident
- b2bcall-fs-config now resolves real directory data: Tenant.telephonyDomain
-> Extension.number, decrypts the password, builds proper directory XML
including a dial-string param (missing it caused originate to fail with
MANDATORY_IE_MISSING instead of the expected USER_NOT_REGISTERED)
- pinned FreeSWITCH's 357737{domain} to a stable value (b2bcall.local) via a
vars.xml patch in the Dockerfile -- it previously used the container's
dynamic IP, which could never match a stored telephony_domain
- added HTTP Basic auth between FreeSWITCH and fs-config
(gateway-credentials, timingSafeEqual comparison) now that the service
returns real secret data, closing the gap flagged as pending in the XML
Curl phase instead of leaving it open
- found and fixed: PermissionGuard's constructor-injected Reflector came
back undefined at runtime under tsx/esbuild (unreliable cross-file
decorator metadata emission) -- fixed with an explicit @Inject(Reflector);
worth watching for in future guards/services run via tsx
- verified end-to-end: create extension -> originate user/<ext> reports
USER_NOT_REGISTERED (found, not registered) -> delete -> back to
SUBSCRIBER_ABSENT (not found); password never reappears in any GET;
unauthenticated fs-config requests get 401
- docs/EXTENSIONS.md
- packages/telephony: TelephonyProvider interface (agente.md secao 25) and
FreeSwitchTelephonyProvider implementation over the 'esl' library
(actively maintained, TypeScript-native, built-in reconnect-with-backoff
satisfying secao 195); normalizeEslEvent() translates raw ESL events into
the internal vocabulary (secao 24)
- apps/freeswitch-events (b2bcall-fs-events): permanent ESL connection,
resubscribes on every reconnect, publishes normalized events to the
'b2bcall:events' Redis pub/sub channel; containerized (Dockerfile +
docker-compose service) since its whole job is reaching the freeswitch
container by internal hostname
- packages/shared: reusable createLogger() (structured JSON per secao 189),
fixed a BigInt serialization crash surfaced by the esl library's error
stats
- found and fixed a real FreeSWITCH 1.11 default: without an explicit
apply-inbound-acl, mod_event_socket silently rejects any non-loopback
connection ('Access Denied, go away.') even with the correct password —
added a dedicated ACL (loopback + the Docker Compose network range, never
0.0.0.0/0) in infrastructure/freeswitch/overrides/autoload_configs/
- verified end-to-end with a local loopback test call: CALL_CREATED ->
CALL_ANSWERED -> CALL_ENDED observed on the Redis channel with the
correct callUuid and hangup cause
- docs/EVENT_SOCKET.md
- POST /auth/login, /auth/refresh, /auth/logout, /auth/select-tenant,
/auth/change-password, GET /auth/tenants — wired to packages/auth
- JwtAuthGuard + DomainExceptionFilter (401/403 without leaking internals)
- LoginRateLimitGuard: Redis-backed 5/min per IP and per email (agente.md
secao 149), safe across multiple API instances
- helmet + restrictive cors (deny-by-default) + global rate limit
- GET /health, /health/live, /health/ready checking Postgres and Redis
- changePassword() added to packages/auth for the mustChangePassword flow
- fixed REDIS_HOST/POSTGRES_HOST docker-compose-only hostnames not
resolving from the host process; added REDIS_URL for host-side use
- verified end-to-end with curl: login, wrong password / unknown email
(same generic error), authenticated route, missing token, refresh
rotation, logout revocation, and the 429 rate limit kicking in after 5
attempts
- packages/auth: Argon2id password hashing, JWT access tokens (jose),
opaque refresh tokens with rotation, generic error messages (no
user-enumeration via timing or message differences)
- roles/permissions/role_permissions/user_roles/sessions/audit_logs schema
(agente.md secoes 142-150); RBAC scope PLATFORM vs TENANT
- withUserContext(): narrow RLS exception so a user can discover their own
tenant_memberships before a tenant is chosen (login flow)
- userHasPermission()/isPlatformUser(): explicit service-layer RBAC checks
(roles/permissions tables are not RLS-protected — documented why in
docs/AUTHENTICATION.md)
- seed: permission catalog, 4 system roles, initial Platform Super Admin
(password written once to FIRST_LOGIN.txt, 600, outside Git)
- automated end-to-end test: login, RBAC check, refresh rotation, logout
- users + tenant_memberships tables (tenant-scoped)
- RLS policy on tenant_memberships using set_config('app.current_tenant_id', ...)
- withTenantContext() helper for transaction-scoped tenant context
- separate non-superuser app role (b2bcall_app): the default Docker postgres
user is SUPERUSER and always bypasses RLS even with FORCE, so the app must
never connect through the migration/owner role. Documented in
docs/TENANT_ISOLATION.md.
- automated isolation test proving tenant A never sees tenant B's data