Usuário reportou que 1502 ligando pra 1503 (dois telefones IP reais,
ambos registrados) não completava a chamada. Três bugs em camadas
diferentes, cada um só confirmado lendo o log/capturando pacote real —
nunca por suposição:
1. Registro: os dois ramais registravam com Contact apontando pro MESMO
IP público compartilhado desta rede (a VM nunca tem IP público
próprio, é sempre RFC1918 atrás do NAT do escritório) — originar uma
chamada tentava mandar o INVITE de volta pra esse IP público, hairpin
NAT clássico, falha instantânea (503). Fix: NDLB-received-in-nat-reg-
contact (Contact salvo vira o IP realmente observado no pacote).
De quebra, aplicado network_mode: host no serviço freeswitch (pedido
explícito do usuário) — tira o Docker NAT/bridge do meio. Quebra em
cascata corrigida: fs-config vira alcançável só via 127.0.0.1:8080
(não mais nome de serviço), workers ESL (fs-events/predictive-dialer)
passam a usar host.docker.internal.
2. Áudio: com o registro corrigido, a chamada completava mas sem RTP —
local-network-acl="localnet.auto" só cobria a subnet da própria
interface do FreeSWITCH, tratando ramais de OUTRAS subnets do mesmo
escritório como "de fora" e trocando o SDP pelo IP público de novo.
Fix: ACL própria (b2bcall_lan, cobre todo RFC1918) referenciada em
local-network-acl.
3. Chamada morrendo sozinha em exatos 32s (Timer H do RFC 3261): mesmo
com áudio ok, o 200 OK que o FreeSWITCH manda pro ramal que recebeu a
chamada ainda tinha Contact com o IP público — o telefone nunca manda
o ACK de volta, FreeSWITCH retransmite sozinho até desistir. Só
confirmado com tcpdump (instalado nesta sessão) capturando o pacote
byte a byte. Fix: ext-rtp-ip/ext-sip-ip (STUN, sempre resolvem pro IP
público) removidos do profile "internal" — sem endereço "externo"
configurado, o FreeSWITCH nunca mais tem como escolher errado.
Confirmado resolvido pelo usuário com chamadas reais nos dois sentidos,
áudio bidirecional, sobrevivendo bem além dos 32s que travavam antes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaBaBSQGhyXGEgtTYZGV8
Pedido do usuário: "pode iniciar a montar o IVR e as rotas de entrada".
Investigando antes de escrever qualquer XML de IVR, achei que NENHUMA
chamada de entrada por tronco tinha como funcionar hoje, IVR ou não:
nenhuma carregava `b2bcall_tenant_id` (só REGISTER de ramal e discagem de
saída setam essa variable), e mesmo corrigindo isso, o profile "external"
apontava pro contexto "public" vanilla — um arquivo ESTÁTICO, que sempre
ganha de uma consulta ao mod_xml_curl, então nunca seria dinâmico
enquanto se chamasse "public".
Perguntei ao usuário a granularidade certa (por DID ou por tronco) antes
de desenhar o schema — escolheu por DID, mais flexível (um tronco pode
carregar vários números com destinos diferentes). `InboundRoute` nova
(RLS real, FORCE ROW LEVEL SECURITY): `didNumber` único GLOBAL entre
tenants (mesma exceção já aceita em Tenant.telephonyDomain) — é a ÚNICA
forma de descobrir de qual tenant é uma chamada de entrada ANTES de
identificar o tenant. Resolvido por fan-out sobre tenants ativos, nunca
uma query sem contexto de RLS.
Dockerfile repontou o profile external pra context="inbound" (sem
arquivo estático, cai no mod_xml_curl de verdade). O XML gerado pra esse
contexto injeta b2bcall_tenant_id + domain_name (achado real: sem setar
domain_name explicitamente, o bridge da "Discagem interna" resolvia pro
domínio GLOBAL default, nunca pro do tenant) e transfere pro dialplan
real do tenant — reaproveita 100% do que já existe, inclusive pickup de
grupo (PHASE 53).
CRUD completo (InboundRoutesController, permissions novas no seed) + tela
"Telefonia > Rotas de Entrada" no frontend.
Testado com uma chamada REAL: um softphone registrado como ramal normal,
outro discando direto pro profile external (porta 5080, sem registrar —
exatamente como um provedor de tronco manda) um DID cadastrado. `show
channels` confirma: tenant certo, contexto certo, domínio certo no
bridge, codec PCMU negociado nos dois lados, ramal tocou e atendeu de
verdade. Detalhes completos, inclusive uma tentativa de teste que falhou
por limitação do canal `loopback` (não um bug), em docs/INBOUND_ROUTES.md.
O IVR em si (menu com play_and_get_digits) fica pra próxima fase — esta
é a fundação sem a qual nada de chamada de entrada funcionava.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaBaBSQGhyXGEgtTYZGV8
Usuário tentou registrar um ramal e não conseguiu — nenhuma porta SIP/RTP
estava publicada no host, só o Event Socket (interno). O teste ponta a
ponta da fase anterior só funcionava porque os softphones de teste
rodavam dentro da mesma rede Docker.
RTP restrito a um range fixo de 200 portas (16384-16584, via sed em
switch.conf.xml) — o default vanilla (~16k portas) é inviável de publicar
uma a uma. docker-compose.yml publica 5060/udp+tcp (SIP) e
16384-16584/udp (RTP); Event Socket continua nunca publicado.
Verificado: iptables -t nat -L DOCKER confirma DNAT correto pras 201
portas; sofia status confirma que o STUN já configurado (external_rtp_ip/
external_sip_ip) resolve pro IP público real da VM, então o SDP vai
anunciar o IP certo — não só o registro, o áudio também deve funcionar.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaBaBSQGhyXGEgtTYZGV8
Pedido explícito do usuário: "testa criando um ramal com callgroup e captura
chamada de outro ramal". O teste com softphones reais (não só leitura de
código) achou 3 problemas que a fase anterior tinha dado como resolvidos:
1. A variable de call group estava com o nome errado (`call-group`, convenção
do Asterisk) e a suposição de que o FreeSWITCH fazia pickup automático só
com ela era falsa. Corrigido pro nome certo (`callgroup`) e pro mecanismo
real (fork de leg `pickup/<grupo>` no bridge + extension `*8` chamando a
application `pickup`, lendo o grupo via `${user_data(...)}`).
2. Implementando o mecanismo acima, achada uma vulnerabilidade real de RCE: o
allowlist de `application` no dialplan nunca bloqueava `${system(...)}`
embutido dentro do `data` de qualquer application já permitida —
`mod_commands` está carregado, então isso era execução de comando
arbitrário no host do FreeSWITCH pra qualquer Tenant Admin. Corrigido com
um segundo allowlist (`ALLOWED_INLINE_API_FUNCTIONS` +
`IsSafeDialplanData`) que só libera funções de leitura seguras
(`user_data`, `escape`, `url_encode`, `url_decode`, `regex`, `strftime`).
3. Registrar um softphone de verdade contra o domínio do tenant (não só curl
no mod_xml_curl) revelava 403 Forbidden: o sofia profile `internal` tinha
`force-register-domain`/`force-subscription-domain`/
`force-register-db-domain` fixados no domínio antigo, ignorando o domínio
de cada tenant. Corrigido no Dockerfile do FreeSWITCH (imagem
reconstruída, não só patch ao vivo).
Testado ponta a ponta com 3 softphones reais (linphone-cli) em containers na
mesma rede Docker: ramal do mesmo grupo captura de verdade uma ligação
tocando em outro ramal via *8 (canais confirmados bridged via `show
channels`); ramal de grupo diferente tenta e falha. RCE confirmado bloqueado
via curl (`${system(id)}` → 400) sem quebrar `${user_data(...)}` legítimo.
TODO.md (PHASE 53) e docs/EXTENSIONS.md atualizados corrigindo as afirmações
incompletas da fase anterior ("nenhuma mudança de infra necessária").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaBaBSQGhyXGEgtTYZGV8
Fecha um risco documentado desde a PHASE 01: os dois processos de host
(fora do Docker) não sobreviviam a um reboot da VM, precisando ser
religados manualmente toda vez. Agora são b2bcall-api.service e
b2bcall-frontend.service (infrastructure/systemd/), enabled, sobrevivem a
reboot como os containers Docker já sobreviviam via restart:
unless-stopped.
Achado real resolvido antes de instalar: sem porta fixa, os dois
disputavam a 3000 por padrão — quem perdesse crashava com EADDRINUSE em
vez de cair pra 3001. Fixado API_PORT=3000 (.env) e PORT=3001
(Environment= direto na unit do frontend — confirmado que colocar em
apps/frontend/.env.local não funciona, o CLI do Next.js decide a porta
antes de aplicar esse arquivo). A ordem de start deixa de importar.
Rodam em modo dev (pnpm dev), não build de produção — decisão deliberada
documentada no README, virar produção é escopo maior.
Testado ponta a ponta: units instaladas e habilitadas, subiram nas portas
certas sem fallback nem crash, journalctl com logs limpos, smoke test de
regressão nas 19 telas do tenant + platform via os processos
supervisionados, todas 200. Achado incidental: as rotas do Next.js dev
levam 10-15s pra compilar na primeira visita — explica os TimeoutError
intermitentes do Puppeteer vistos ao longo da sessão, não é um bug.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BFaBaBSQGhyXGEgtTYZGV8
- Investigated the real callcenter_config command surface via
'help callcenter_config' on the running FreeSWITCH before writing any
code: queues only have load/unload/reload (static XML + reload, no
'queue add' exists), while agents and tiers are fully dynamic via ESL
commands (agent add, tier add) -- no file involved. This shapes the next
phase (Agents/Tiers) differently from this one.
- queues table (tenant-scoped, RLS): strategy, moh/announce, wait times,
tier rules, discard/abandoned handling, skip-agents-with-external-calls,
recording_enabled
- packages/telephony: buildQueueXml()
- infrastructure/freeswitch: our own callcenter.conf.xml override (empties
the vanilla static agents/tiers -- those become fully dynamic in the next
phase) that includes callcenter_queues.conf.d/*.xml via X-PRE-PROCESS,
same pattern as the Sofia gateway directory
- apps/api/src/queues: CRUD (POST/GET/GET:id/DELETE) using the queues.view/
.manage permissions already in the seed
- b2bcall-fs-config (queue-sync.ts): one XML file per queue on a shared
volume, synced via Redis pub/sub (b2bcall:queues:sync) on create/delete
and once at boot -- same shape as trunk-sync.ts
- confirmed manually against the real FreeSWITCH, before coding the sync
logic: 'queue load <name>' fails ('Invalid Queue not found!') for a file
added after boot -- needs 'reloadxml' first to repopulate the in-memory
XML tree from disk; after that, 'queue reload <name>' alone handles both
create and update, no need to distinguish load vs reload
- verified end-to-end: created a queue (ROUND_ROBIN, maxWaitTime=120,
discardAbandonedAfter=90), 'callcenter_config queue list' showed the
correct values on the FreeSWITCH side; deleted it, list went back empty
docs/QUEUES.md
- extensions table (tenant-scoped, RLS): number, sip_password_enc
(AES-256-GCM via packages/shared/src/crypto.ts), caller_id, context,
sofia_profile, codecs, max_registrations
- apps/api/src/extensions: CRUD (POST/GET/GET:id/DELETE), protected by a
new generic PermissionGuard (@RequirePermission decorator), tenant
resolved only from the JWT (never trusted from the client)
- SIP password is returned in plaintext only once, in the create response;
toPublicExtension() explicitly destructures the encrypted field out
(not a spread) so it can't leak by accident
- b2bcall-fs-config now resolves real directory data: Tenant.telephonyDomain
-> Extension.number, decrypts the password, builds proper directory XML
including a dial-string param (missing it caused originate to fail with
MANDATORY_IE_MISSING instead of the expected USER_NOT_REGISTERED)
- pinned FreeSWITCH's 357737{domain} to a stable value (b2bcall.local) via a
vars.xml patch in the Dockerfile -- it previously used the container's
dynamic IP, which could never match a stored telephony_domain
- added HTTP Basic auth between FreeSWITCH and fs-config
(gateway-credentials, timingSafeEqual comparison) now that the service
returns real secret data, closing the gap flagged as pending in the XML
Curl phase instead of leaving it open
- found and fixed: PermissionGuard's constructor-injected Reflector came
back undefined at runtime under tsx/esbuild (unreliable cross-file
decorator metadata emission) -- fixed with an explicit @Inject(Reflector);
worth watching for in future guards/services run via tsx
- verified end-to-end: create extension -> originate user/<ext> reports
USER_NOT_REGISTERED (found, not registered) -> delete -> back to
SUBSCRIBER_ABSENT (not found); password never reappears in any GET;
unauthenticated fs-config requests get 401
- docs/EXTENSIONS.md
- apps/freeswitch-config (b2bcall-fs-config): Fastify service implementing
the mod_xml_curl HTTP protocol (form-encoded POST -> XML response),
containerized, no host port published
- reactivated mod_xml_curl in FreeSWITCH, binding restricted to
directory|dialplan only (configuration was removed after testing showed
it firing several unnecessary HTTP round-trips at boot for module
configs we don't need dynamic — matches agente.md's own 'don't put every
critical config through XML Curl' guidance)
- no extensions/dialplan tables exist yet (next phases), so the service
always answers 'not found' for now — this phase only proves the wire
protocol works without breaking the static vanilla config fallback
- verified end-to-end: user/8888 (nowhere) -> SUBSCRIBER_ABSENT via
fs-config; user/1000 (static vanilla extension) -> USER_NOT_REGISTERED,
proving FreeSWITCH correctly falls through to static XML when xml_curl
says not found
- docs/XML_CURL.md, including the not-yet-authenticated endpoint note (fine
while it only returns not-found; needs gateway-credentials before serving
real directory/dialplan data)
- packages/telephony: TelephonyProvider interface (agente.md secao 25) and
FreeSwitchTelephonyProvider implementation over the 'esl' library
(actively maintained, TypeScript-native, built-in reconnect-with-backoff
satisfying secao 195); normalizeEslEvent() translates raw ESL events into
the internal vocabulary (secao 24)
- apps/freeswitch-events (b2bcall-fs-events): permanent ESL connection,
resubscribes on every reconnect, publishes normalized events to the
'b2bcall:events' Redis pub/sub channel; containerized (Dockerfile +
docker-compose service) since its whole job is reaching the freeswitch
container by internal hostname
- packages/shared: reusable createLogger() (structured JSON per secao 189),
fixed a BigInt serialization crash surfaced by the esl library's error
stats
- found and fixed a real FreeSWITCH 1.11 default: without an explicit
apply-inbound-acl, mod_event_socket silently rejects any non-loopback
connection ('Access Denied, go away.') even with the correct password —
added a dedicated ACL (loopback + the Docker Compose network range, never
0.0.0.0/0) in infrastructure/freeswitch/overrides/autoload_configs/
- verified end-to-end with a local loopback test call: CALL_CREATED ->
CALL_ANSWERED -> CALL_ENDED observed on the Redis channel with the
correct callUuid and hangup cause
- docs/EVENT_SOCKET.md
- infrastructure/freeswitch/Dockerfile: debian:trixie-slim + SignalWire
packaged freeswitch-meta-vanilla, avoiding a C/C++ build on a 1.9GB RAM VM
- FREESWITCH_PAT used only via Docker BuildKit secret, apt credentials file
created and deleted within the same RUN — verified absent from the final
image with docker history
- minimal module set (agente.md secao 15): sofia, event_socket, commands,
dptools, callcenter, avmd, curl, local_stream, etc. mod_xml_curl installed
but disabled — it refuses to load without a configured gateway-url, which
will exist once b2bcall-fs-config is built
- entrypoint.sh rotates the Event Socket password away from the 'ClueCon'
default at container runtime (never baked into the image); fails loudly if
ESL_PASSWORD is unset
- port 8021 not published to the host; only reachable from other containers
on the compose network
- found and fixed: freeswitch-conf-vanilla is a Recommends (not a Depends)
of freeswitch-meta-vanilla, so --no-install-recommends silently produced
an empty /etc/freeswitch and a crash loop
- verified end-to-end: fs_cli status via ESL with the custom password,
default password rejected, expected modules loaded, healthcheck green,
~44MB RAM usage
- docs/FREESWITCH.md, docs/NETWORK_ARCHITECTURE.md (network_mode decision
deferred until a real SIP trunk exists)