ca49504f0a312a69d5451d91934a2ce62ce9b46c
8 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
| cb6d343b2e |
feat(dialer): CPS Limiter + Predictive Dialer Engine
Fecha agente.md secao 72-86 (motor preditivo) e 77-79 (CPS distribuido,
reserva de leads, lock de campanha). Uma campanha RUNNING agora origina
chamadas sozinha, respeitando capacidade de agentes, CPS hierarquico e
taxa de abandono — sem intervencao manual.
Deliberadamente fora do escopo (agente.md secao 72: "nao e' so' `for lead
-> originate`"): mod_avmd (opcional), callbacks agendados, disposicoes de
agente — ficam pra fase CDR.
## Novo servico apps/predictive-dialer
Mesmo padrao arquitetural de fs-events/fs-config: Node standalone em
Docker, ESL propria, tick a cada 2s sobre tenants ativos x campanhas
RUNNING/WAITING_SCHEDULE.
- Lock de campanha (dialer:campaign:{id}, secao 79): TTL/ownership/
renewal/safe-release via Lua compare-and-delete.
- CPS distribuido (secao 77, 62): token bucket janela 1s, hierarquia
GLOBAL/TENANT/TRUNK/CAMPAIGN numa unica chamada Lua atomica — nivel
esgotado bloqueia todos SEM incremento parcial dos que passariam.
- Reserva atomica de leads (secao 78): FOR UPDATE SKIP LOCKED dentro da
mesma transacao withTenantContext.
- CallAttempt/CampaignStats (schema novo): state machine da chamada
(secao 82) + EWMA (secao 75) de answer_probability/average_answer_delay/
average_talk_time/abandon_rate por campanha.
- Capacidade em tempo real + pacing (secao 73-76, 84-85): conta agentes
por estado via Tier->Agent.state, previsao de liberacao (horizonte
unico de 15s, simplificacao documentada dos 4 buckets da especificacao),
controle de abandono reduz pacing progressivamente, nunca origina sem
capacidade prevista.
## Modo simulacao (secao 185-186)
DIALER_SIMULATION=true (default, ja estava no .env desde o inicio da
sessao) sorteia ANSWER/BUSY/NO_ANSWER/FAILED em software, sem PSTN real.
So' quando ANSWERED e' que uma chamada sintetica (null/dummy, sem PSTN)
entra na fila real via mod_callcenter de verdade — escolha deliberada pra
maximizar codigo real exercitado em vez de simular tudo em memoria. Os
identificadores da secao 81 (b2bcall_tenant_id/call_id/attempt_id/
campaign_id/lead_id) vao como channel variables nessa perna, entregando
tenantId real no WebSocket sem fan-out.
Real Outbound Safety (secao 186): as duas flags checadas no boot, nunca
ativadas automaticamente — caminho PSTN real implementado mas nunca
exercitado (sem trunk/operadora real neste laboratorio).
## Dois bugs reais achados e corrigidos testando esta fase
- Perna sintetica (null/dummy) nao tem midia do outro lado — nunca
desligava sozinha depois de bridgear com um agente. Corrigido com
hangup agendado via uuid_kill no talk_time simulado.
- Corrida entre queue:sync e tier:sync (dois canais Redis independentes,
sem ordem garantida): atribuir tier logo depois de criar a fila podia
rodar tier add antes do queue reload terminar ("-ERR Queue not found!",
erro real, diferente do ja conhecido "already exist"). Corrigido com
retry curto (ate 3 tentativas) em agent-sync.ts::addTierWithRetry.
## GET /campaigns/:id/stats
Secao 227.7 "visualizar pacing" — CampaignStats + agentes por estado +
calls em andamento, sem esperar a fase Frontend.
Verificado ponta a ponta: campanha RUNNING originando 3 tentativas por
tick, outcomes simulados corretos com retry agendado (BUSY 15min/
NO_ANSWER 60min/FAILED 30min), uma tentativa ANSWERED completando o ciclo
real inteiro (fila -> agente -> bridge -> hangup -> EWMA atualizada),
stop nao derruba chamada ativa (secao 66), calls_answered=3 confirmado no
`queue list` do FreeSWITCH. CPS limiter e lock de campanha testados
isoladamente (hierarquia sem incremento parcial, ownership nunca
roubado). typecheck do workspace inteiro limpo.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1HxY46WGU4G1zmVDNKcWw
|
|||
| b940ce3e63 |
fix(events): eventos CUSTOM do ESL nunca eram entregues
Achado ao iniciar a fase de Realtime Monitoring: nenhum evento CUSTOM (sofia::register, sofia::gateway_state, callcenter::info) jamais chegou em b2bcall-fs-events nesta sessao, apesar do estado real do FreeSWITCH mudar de verdade (confirmado originando uma chamada de teste pra dentro de uma fila real com agente logado). Isso deixava dois gaps documentados como "nao verificado" em docs/TRUNKS.md e docs/AGENTS.md. Causa raiz: `event_json(...SUBSCRIBED_EVENTS)` mandava "CUSTOM" como ultimo token do comando `event json`, sem nenhum subclass depois. O mod_event_socket do FreeSWITCH exige que os subclasses (callcenter::info, sofia::register, ...) venham imediatamente depois do token CUSTOM no mesmo comando — sem isso, zero eventos CUSTOM sao entregues, de qualquer subclass. Corrigido separando PLAIN_EVENTS (nomes normais, cada um vira um listener .on()) de CUSTOM_SUBCLASSES (so compoe o comando de subscricao — o client ESL sempre emite "CUSTOM" como nome de evento, com o subclass real no header Event-Subclass). Corrigido tambem um bug de nome de campo: normalizeCustomEvent lia CC-Agent-Status, que nao existe; o campo real e CC-Agent-State. Com o pipeline corrigido, chegam eventos ricos de callcenter::info nunca antes observados: agent-offering, bridge-agent-fail, members-count (fila em tempo real) e member-queue-end (com CC-Cause/CC-Cancel-Reason e timestamps de entrada/saida — atendida vs. abandonada). Adicionados como novos tipos normalizados: AGENT_OFFERED_CALL, AGENT_BRIDGE_FAILED, QUEUE_MEMBER_COUNT, QUEUE_MEMBER_LEFT. De quebra, achado e corrigido um segundo bug real ao reverificar Trunks com o pipeline de eventos funcionando: `sofia profile external rescan` nunca descarregava um gateway cujo arquivo .xml foi apagado (fica fantasma na memoria do Sofia indefinidamente). trunk-sync.ts agora roda `sofia profile external killgw <nome>` pra cada gateway removido, antes do rescan. Reverificado ponta a ponta pra ambos os bugs: - Trunk com register:true apontando pra host inexistente: Trunk.status no banco passa de UNKNOWN pra FAILED sozinho, via evento, sem polling. - Trunk apagado via API: gateway some imediatamente de `sofia status gateway`, sem esperar reinicio de profile. - Chamada de teste real numa fila com agente logado: members-count, agent-offering, agent-state-change (CC-Agent-State correto: Waiting/ Receiving), bridge-agent-fail, todos chegando certos no canal Redis b2bcall:events. typecheck do workspace inteiro limpo. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X1HxY46WGU4G1zmVDNKcWw |
|||
| a05104a05f |
feat(agents): agentes, tiers e pausas — call center completo
Fecha agente.md secao 45-49/52. Depois desta fase, um usuario autenticado consegue logar como agente, entrar numa fila real, se pausar e voltar, tudo refletido de verdade no FreeSWITCH. Schema (migration 20260828124245_agents): - agents (tenant-scoped, RLS): User -> Extension -> identidade de agente, state (enum AgentState de 8 valores) espelhando o estado real, so alterado via login/logout/pause/resume, nunca escrito direto pela API. - tiers: Queue<->Agent (level/position 1:1 com mod_callcenter). - agent_sessions: um ciclo login->logout por linha. - agent_state_events: historico de transicoes de estado. - pause_reasons / agent_pause_events (secao 48). Dois bugs reais corrigidos em FreeSwitchTelephonyProvider, presentes desde a fase de Event Socket original: - "queue add/del member" nao existe no mod_callcenter — membership de fila usa tier add/tier del. So foi pego agora ao confirmar de novo a sintaxe via `help callcenter_config` antes de codar esta fase. - addAgent/removeAgent nao existiam ainda (agent add/del). Mecanismo de sync: agentes e tiers nao tem representacao em XML, so comando ESL direto — diferente do padrao "regenera todos os arquivos" usado em Trunks/Queues. apps/api publica uma mensagem por acao com payload (b2bcall:agents:sync, b2bcall:tiers:sync); b2bcall-fs-config aplica o comando correspondente (agent-sync.ts). Achados confirmados manualmente contra o FreeSWITCH real antes de codar: - `agent add`/`tier add` nao sao idempotentes (erro em duplicata) — sync ignora esse erro (.catch), condicao esperada em resync. - `agent del`/`tier del` em algo inexistente nao da erro — seguro chamar sem checar existencia antes. - `agent set status` so aceita 3 valores exatos (Available/On Break/ Logged Out) — testado deliberadamente com valor invalido. - Corrida real: atribuir tier antes do primeiro login do agente falha silenciosamente do lado do FreeSWITCH (agente so existe la a partir do `agent add` no login). Login sempre re-sincroniza todos os tiers do agente depois de garantir que ele existe — auto-correcao confirmada no teste ponta a ponta. apps/api: AgentsController (CRUD), AgentsMeController (login/logout/ pause/resume — sempre resolve o agente via JWT, nunca um agentId arbitrario do client), PauseReasonsController (CRUD), QueueAgentsController (POST/DELETE de tier em /queues/:id/agents). Verificado ponta a ponta via curl + fs_cli contra o FreeSWITCH real: login -> Available, pause -> On Break, resume -> Available, logout -> Logged Out, todos batendo entre Agent.state (banco) e `agent list` (FreeSWITCH). typecheck do workspace inteiro limpo. ~350MB de memoria total (docker stats). Documentado em docs/AGENTS.md, incluindo lacuna conhecida: estados derivados de chamada (RINGING/IN_CALL/WRAP_UP/RESERVED) dependem do evento CUSTOM callcenter::info, ainda nao comprovado chegando em fs-events nesta sessao (mesma lacuna de sofia::gateway_state ja documentada em docs/TRUNKS.md) — precisa de uma chamada real passando pela fila pra investigar. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X1HxY46WGU4G1zmVDNKcWw |
|||
| 6628578c42 |
feat: implement Call Center queues (mod_callcenter)
- Investigated the real callcenter_config command surface via
'help callcenter_config' on the running FreeSWITCH before writing any
code: queues only have load/unload/reload (static XML + reload, no
'queue add' exists), while agents and tiers are fully dynamic via ESL
commands (agent add, tier add) -- no file involved. This shapes the next
phase (Agents/Tiers) differently from this one.
- queues table (tenant-scoped, RLS): strategy, moh/announce, wait times,
tier rules, discard/abandoned handling, skip-agents-with-external-calls,
recording_enabled
- packages/telephony: buildQueueXml()
- infrastructure/freeswitch: our own callcenter.conf.xml override (empties
the vanilla static agents/tiers -- those become fully dynamic in the next
phase) that includes callcenter_queues.conf.d/*.xml via X-PRE-PROCESS,
same pattern as the Sofia gateway directory
- apps/api/src/queues: CRUD (POST/GET/GET:id/DELETE) using the queues.view/
.manage permissions already in the seed
- b2bcall-fs-config (queue-sync.ts): one XML file per queue on a shared
volume, synced via Redis pub/sub (b2bcall:queues:sync) on create/delete
and once at boot -- same shape as trunk-sync.ts
- confirmed manually against the real FreeSWITCH, before coding the sync
logic: 'queue load <name>' fails ('Invalid Queue not found!') for a file
added after boot -- needs 'reloadxml' first to repopulate the in-memory
XML tree from disk; after that, 'queue reload <name>' alone handles both
create and update, no need to distinguish load vs reload
- verified end-to-end: created a queue (ROUND_ROBIN, maxWaitTime=120,
discardAbandonedAfter=90), 'callcenter_config queue list' showed the
correct values on the FreeSWITCH side; deleted it, list went back empty
docs/QUEUES.md
|
|||
| 0720a0efe3 |
feat: implement Dialplan with structured editor and versioning
- dialplan_extensions table (tenant-scoped, RLS): structured editor per agente.md secao 43 -- context, condition field/expr, actions/anti-actions (JSON), continue, order, enabled. One condition per extension (deliberate simplification vs raw FreeSWITCH's multi-condition extensions). - dialplan_versions table (tenant-scoped, RLS): generate/validate/version/ activate flow (secao 44). Reactivating an older version IS the rollback mechanism -- no separate endpoint needed. - apps/api/src/dialplan: extensions CRUD + versions/generate (builds XML, validates well-formedness with fast-xml-parser, saves as DRAFT) + versions/:id/activate (atomically flips ACTIVE, supersedes the previous one). Reused freeswitch.view/.configure permissions rather than inventing new ones not in the agente.md permission list. - packages/telephony: buildDialplanXml() plus ALLOWED_DIALPLAN_APPLICATIONS, an explicit allowlist (answer/bridge/playback/hangup/set/export/... -- deliberately no system/exec/socket) guarding against a tenant configuring a dialplan action that runs arbitrary commands on the FreeSWITCH host (agente.md secao 180) - b2bcall-fs-config resolves dialplan dynamically per call (unlike Trunks' file+rescan approach -- dialplan is fetched fresh via mod_xml_curl on every call anyway) by tenant id from the variable_b2bcall_tenant_id channel variable already injected at directory resolution, then serving whichever DialplanVersion is ACTIVE for that context - verified end-to-end: created a rule for destination_number 7000, generated and activated v1, originated a call that actually routed through the dialplan (not bypassing it via &app()) -- CALL_CREATED -> CALL_ANSWERED -> CALL_ENDED with the correct tenantId throughout. Created and activated a v2, then rolled back to v1 by reactivating it; status transitions (ACTIVE/SUPERSEDED) all confirmed via the API. CRITICAL FINDING, fixed in this same phase: deliberately testing that the application allowlist rejects 'system' got back 201 instead of 400 -- NestJS's ValidationPipe had been silently inert across all of apps/api's @Body() DTOs since the API was first created. Root cause: running via (esbuild) instead of a real build -- esbuild doesn't always resolve cross-file parameter types for design:paramtypes metadata, and Nest skips validation without any error when it can't determine the DTO class. Fixed by always building with tsc before running (tsc && tsx dist/main.js -- still via tsx because internal workspace packages aren't built to JS yet). Re-verified with two deliberate bad-input tests post-fix, both correctly rejected with 400. A stray malicious test row (dialplan action 'system') created while the bug was live was deleted; it was never baked into an activated version, so nothing could have executed it. See docs/VALIDATION_PIPE_BUG.md for the full writeup. docs/DIALPLAN.md, docs/VALIDATION_PIPE_BUG.md, docs/EXTENSIONS.md updated |
|||
| 4c638ad496 |
feat: implement Trunks with real FreeSWITCH gateway sync
- trunks table (tenant-scoped, RLS): host/proxy/realm, register, username/password_enc (AES-256-GCM), dtmf_mode, ping, transport, and a status/status_updated_at pair meant to be driven by FreeSWITCH events - apps/api/src/trunks: CRUD (POST/GET/GET:id/DELETE), same RBAC/tenant pattern as Extensions, password never exposed in any GET - packages/telephony: buildGatewayXml() generates a Sofia gateway XML file - b2bcall-fs-config now writes sip_profiles/external/<trunk_id>.xml (shared Docker volume with FreeSWITCH -- the vanilla external profile already includes external/*.xml) and runs 'sofia profile external rescan' over ESL; syncs on boot and on demand via Redis pub/sub (b2bcall:trunks:sync), since apps/api runs on the host and fs-config has no port published to reach directly - added FreeSwitchTelephonyProvider.waitUntilConnected() to fix a startup race: the first sync ran before the ESL connection had settled, logging a harmless but noisy error - verified end-to-end with a fake host: create trunk -> gateway file written -> FreeSWITCH shows the real gateway (FAIL_WAIT, expected) -> delete -> file removed (cleanup also correctly swept the stale 'example.com' gateway that had been copied into the volume from the vanilla image) - apps/freeswitch-events/src/trunk-status.ts: written to update Trunk.status from sofia::gateway_state events, using the same normalizeEslEvent path already proven for CHANNEL_* events - KNOWN GAP, documented rather than glossed over: monitored fs-events for ~90s while the gateway visibly transitioned states in FreeSWITCH (FAIL_WAIT/DOWN) and no sofia::gateway_state event was observed arriving. CUSTOM/sofia::* events have not actually been proven working end-to-end in this session -- only CHANNEL_* events have been. Needs verification against a real SIP target before the status auto-update can be trusted in production. See docs/TRUNKS.md and TODO.md. - docs/TRUNKS.md |
|||
| c03c6d4eaa |
feat: implement Extensions with real FreeSWITCH directory integration
- extensions table (tenant-scoped, RLS): number, sip_password_enc
(AES-256-GCM via packages/shared/src/crypto.ts), caller_id, context,
sofia_profile, codecs, max_registrations
- apps/api/src/extensions: CRUD (POST/GET/GET:id/DELETE), protected by a
new generic PermissionGuard (@RequirePermission decorator), tenant
resolved only from the JWT (never trusted from the client)
- SIP password is returned in plaintext only once, in the create response;
toPublicExtension() explicitly destructures the encrypted field out
(not a spread) so it can't leak by accident
- b2bcall-fs-config now resolves real directory data: Tenant.telephonyDomain
-> Extension.number, decrypts the password, builds proper directory XML
including a dial-string param (missing it caused originate to fail with
MANDATORY_IE_MISSING instead of the expected USER_NOT_REGISTERED)
- pinned FreeSWITCH's 357737{domain} to a stable value (b2bcall.local) via a
vars.xml patch in the Dockerfile -- it previously used the container's
dynamic IP, which could never match a stored telephony_domain
- added HTTP Basic auth between FreeSWITCH and fs-config
(gateway-credentials, timingSafeEqual comparison) now that the service
returns real secret data, closing the gap flagged as pending in the XML
Curl phase instead of leaving it open
- found and fixed: PermissionGuard's constructor-injected Reflector came
back undefined at runtime under tsx/esbuild (unreliable cross-file
decorator metadata emission) -- fixed with an explicit @Inject(Reflector);
worth watching for in future guards/services run via tsx
- verified end-to-end: create extension -> originate user/<ext> reports
USER_NOT_REGISTERED (found, not registered) -> delete -> back to
SUBSCRIBER_ABSENT (not found); password never reappears in any GET;
unauthenticated fs-config requests get 401
- docs/EXTENSIONS.md
|
|||
| d2ea83c06a |
feat: activate mod_xml_curl with b2bcall-fs-config (agente.md secao 26)
- apps/freeswitch-config (b2bcall-fs-config): Fastify service implementing the mod_xml_curl HTTP protocol (form-encoded POST -> XML response), containerized, no host port published - reactivated mod_xml_curl in FreeSWITCH, binding restricted to directory|dialplan only (configuration was removed after testing showed it firing several unnecessary HTTP round-trips at boot for module configs we don't need dynamic — matches agente.md's own 'don't put every critical config through XML Curl' guidance) - no extensions/dialplan tables exist yet (next phases), so the service always answers 'not found' for now — this phase only proves the wire protocol works without breaking the static vanilla config fallback - verified end-to-end: user/8888 (nowhere) -> SUBSCRIBER_ABSENT via fs-config; user/1000 (static vanilla extension) -> USER_NOT_REGISTERED, proving FreeSWITCH correctly falls through to static XML when xml_curl says not found - docs/XML_CURL.md, including the not-yet-authenticated endpoint note (fine while it only returns not-found; needs gateway-credentials before serving real directory/dialplan data) |