Member Junction
    Preparing search index...

    Document / NoSQL family of the EDS-consuming ingestion connector (MongoDB today). The heart's ingestion flow applies unchanged — the MongoDB EDS driver's RunView accepts the same structured incrementalSince bound + ordering columns as the SQL drivers (rendering them to a Mongo query, coercing the watermark to a Date, itself) — so this family supplies only the one document-specific difference:

    • Non-authoritative discovery — document introspection SAMPLES documents (a bounded scan), so it does NOT enumerate the full field gamut; a field absent from a refresh proves nothing and must NOT drive deactivation. DiscoveryIsAuthoritative stays false (§7).

    Incremental sync on a document store depends on the collection carrying a conventional "last changed" field (name-detected by the heart); a collection without one syncs full, not incremental.

    Still abstract: a thin per-engine leaf (MongoConnector) registers it via @RegisterClass.

    Hierarchy (View Summary)

    Index

    Constructors

    Properties

    ExternalDataSourceIDKey: "externalDataSourceID" = 'externalDataSourceID'

    Configuration key on CompanyIntegration.Configuration naming the shared EDS data-source row.

    Accessors

    • get FetchChangesTimeoutMs(): number

      Per-connector override for the FetchChanges timeout, in milliseconds. null → use DEFAULT_OPERATION_TIMEOUTS.FetchChangesMs (30s).

      Raise this when a single page is legitimately slow: a connector that fans out one request per parent (ORCID's per-iD /record, any second-layer object) does N requests inside ONE FetchChanges call, so its page time scales with BatchSize — and with how much concurrency the engine's adaptive controller currently allows. Under the fixed 30s, a page that comfortably fit when parallel no longer fit once the controller had cut concurrency, which forced connector authors to shrink BatchSize for the sequential worst case and waste the parallel headroom the rest of the time.

      Note the timeout does NOT itself cut concurrency: ClassifyError gives it NETWORK_TIMEOUT, and only RATE_LIMIT_EXCEEDED feeds the adaptive limiter. A timing-out page simply never ramps the rate UP (the ramp needs a clean fetch), and the object ends incomplete — reported as FETCH_ABORTED_INCOMPLETE. Earlier revisions of this comment described a self-reinforcing timeout→concurrency-cut spiral; that is not what the code does.

      Precedence (highest first): CompanyIntegration.Configuration.fetchTimeoutMs → this property → DEFAULT_OPERATION_TIMEOUTS.FetchChangesMs. Deployments therefore keep the last word without a code change, while a connector that KNOWS it is slow ships a sane default.

      Returns number

    • get IntegrationName(): string

      The canonical integration name (e.g., "HubSpot", "Rasa.io"). Used by GetActionGeneratorConfig() and IntegrationActionExecutor to match connectors to action Config.IntegrationName.

      Override in subclasses. Defaults to the class name.

      Returns string

    • get MaxConcurrencyHint(): number

      Highest SAFE per-layer concurrency the source tolerates (plan.md §7 peak parallelization) — the ceiling the engine's adaptive controller ramps toward. null → use configured syncConcurrency.

      Returns number

    • get MonotonicWatermark(): boolean

      We fetch strictly in watermark order, so the last batch's max watermark IS the true high-water mark and an updated row always re-surfaces at a new, higher watermark — the engine can safely narrow the next incremental to it (§ MonotonicWatermark).

      Returns boolean

    • get RateLimitPolicy(): RateLimitPolicy

      Token-bucket rate-limit policy for this connector's source API (plan.md §7 peak-aware rate limiting). null → the engine derives a conservative rate from Integration.BatchRequestWaitTime. Override to push to the source's real limits.

      Returns RateLimitPolicy

    • get SupportsBatchWrite(): boolean

      Whether this connector supports batched target writes (plan.md §7 aggressive batching).

      Returns boolean

    • get SupportsCreate(): boolean

      Whether this connector supports creating new records in the external system.

      Returns boolean

    • get SupportsDelete(): boolean

      Whether this connector supports deleting records from the external system.

      Returns boolean

    • get SupportsGet(): boolean

      Whether this connector supports reading/fetching records. Always true.

      Returns boolean

    • get SupportsListing(): boolean

      Whether this connector supports paginated listing of records.

      Returns boolean

    • get SupportsSearch(): boolean

      Whether this connector supports searching/querying records with filters.

      Returns boolean

    • get SupportsUpdate(): boolean

      Whether this connector supports updating existing records in the external system.

      Returns boolean

    • get SupportsUpsert(): boolean

      Whether this connector supports idempotent upserts (create-or-update keyed by a unique business property). Connectors override this AND Upsert to enable it.

      Returns boolean

    Methods

    • Builds a CRUDResult for a record CREATE, failing LOUDLY when the external system returned no usable record ID. A 2xx response with an empty/undefined ID means the create did not durably produce a record we can track — returning Success:true there silently loses the record and causes duplicate creates on the next sync (the HubSpot-association class of bug, fixed in next commit 9f718a7e). This makes that failure explicit at the connector boundary.

      Parameters

      • externalID: string
      • statusCode: number
      • objectName: string

      Returns CRUDResult

    • Coerce a raw watermark value into a Date for ExternalRecord.ModifiedAt (undefined when unparseable). Delegates full ISO-8601 date-times to the shared parseIso8601AsUtc helper — the same one the EDS drivers use for incremental-watermark literal formatting and Mongo date coercion — so a ZONELESS ISO string is interpreted as UTC consistently everywhere, not via a locally-reimplemented regex. Falls back to Date's own parsing for non-ISO shapes (e.g. a date-only string) that helper intentionally rejects. (In practice the driver returns Dates from RunView; this string path is the fallback.)

      Parameters

      • value: unknown

      Returns Date

    • Deterministic content key for a genuinely PK-less row (rare); lets such tables still dedupe. Excludes watermarkField when given — the watermark column changes on every update by definition, so including it would mint a new ExternalID (and therefore a new inserted record) on every update to the same PK-less row instead of matching its prior identity.

      Parameters

      • row: Record<string, unknown>
      • OptionalwatermarkField: string

      Returns string

    • Discovery via the connector's READ PATH (FetchChanges), TIME-BOUNDED — the way to gather a statistically-significant sample when a single DiscoverFields sample is too small to PROVE a key. "Discovery is the sync read path with the save removed."

      Loops FetchChanges as a read-only FULL fetch (WatermarkValue=null, nothing persisted), threading pagination/keyset cursors across batches, and streams every record through the data-informed field + provable-PK inference. It stops at the discovery TIME BUDGET (default 5 min), or a record cap, or source exhaustion — whichever comes first — so the provable-PK decision is made on as much real data as the budget allows. It NEVER fabricates a key: if even this larger sample yields no provable single/composite PK, the field set comes back PK-less and the object is honestly not added.

      Falls back to the single-sample DiscoverFields if the read path can't run for this object (e.g. a connector whose FetchChanges needs an already-persisted IO row that doesn't exist yet).

      Parameters

      • companyIntegration: MJCompanyIntegrationEntity
      • objectName: string
      • contextUser: UserInfo
      • Optionalopts: {
            BatchSize?: number;
            MaxRecords?: number;
            OnFallback?: (err: unknown) => void;
            TimeBudgetMs?: number;
        }
        • OptionalBatchSize?: number
        • OptionalMaxRecords?: number
        • OptionalOnFallback?: (err: unknown) => void

          Called when streaming fails and this method degrades to single-sample DiscoverFields. The degradation is silent otherwise, and it is not a small one: the fallback returns the catalog's own description, which carries no observed widths — so the caller believes it sampled, and the object keeps whatever width the catalog guessed.

        • OptionalTimeBudgetMs?: number

      Returns Promise<ExternalFieldSchema[]>

    • Stage-2 field discovery for sources WITHOUT a describe/introspection endpoint (file feeds, undocumented JSON list endpoints): stream the source's actual records — READ-ONLY, no save, no ack — and derive the full field set + data-informed PK/uniqueness/nullability from the gathered statistics. The connector supplies whatever read-only fetch yields the records; this helper turns that stream into ExternalFieldSchema[].

      Why data-informed: streaming the real values lets pickPrimaryKeyFromStats pick the PK from evidence (uniqueness/non-null statistics) COMBINED with the naming convention, rather than a name guess alone. The PK is a SOFT key, so the pick is best-available, not strict-significance: a confident unique+non-null column wins outright; otherwise a near-unique / convention-named column is taken as a soft key (a PK-less object would stall CodeGen). The scan is time-bounded — it stops on exhaustion OR opts.Discovery.TimeBudgetMs; more rows simply mean stronger claims.

      Provable-only encoding into the standard flags:

      • IsPrimaryKey — set ONLY on the single statistics-first pick. Multiple equally-ranked unique columns leave PK unset here (ambiguous → the pipeline's SoftPKClassifier LLM tiebreaker decides, fed these same stats). Zero unique columns → no PK is fabricated.
      • IsUniqueKey — set when the column was all-distinct over the scan AND uniqueness was provable (the distinct-cap wasn't hit).
      • AllowsNull — asserted true ONLY when a null/absent value was actually observed; otherwise left undefined (permissive default). Never fabricates NOT NULL — critical under a time-capped partial scan where unseen rows could still be null.

      The PK emitted here is SOFT (it rides additionalSchemaInfo via the persist + DDL path; it is NEVER a hard DB key), so a wrong inference can never reject a valid row — the engine dedupes via the record-map. IsReadOnly defaults to true (stream discovery targets read feeds); a writable source overrides via opts.ReadOnly.

      Parameters

      • records:
            | AsyncIterable<Record<string, unknown>, any, any>
            | Iterable<Record<string, unknown>, any, any>

        A read-only sync/async iterable of source records (the caller's fetch yields them).

      • Optionalopts: { Discovery?: StreamDiscoveryOptions; Pk?: PkPickOptions; ReadOnly?: boolean }

      Returns Promise<ExternalFieldSchema[]>

    • The record source DiscoverFieldsViaFetch streams for field/PK inference. Default: loop FetchChanges (full fetch), yielding each record's fields until maxRecords. A protocol subclass (e.g. REST) overrides this to sample a template-var CHILD with the correct record-constrained, recursive stream. Yields plain field maps; the caller stops it at maxRecords.

      Parameters

      • companyIntegration: MJCompanyIntegrationEntity
      • objectName: string
      • contextUser: UserInfo
      • batchSize: number
      • maxRecords: number
      • OptionaldeadlineMs: number

        Epoch ms this sample must not outlive, forwarded to the connector as FetchContext.DeadlineMs. Optional so existing overrides of this method keep compiling and keep their current behaviour.

      • OptionalwatchKey: string

        Watchdog key for this sample, when the caller registered one. Diagnostics only.

      Returns AsyncGenerator<Record<string, unknown>>

    • Parse a Retry-After / rate-limit signal out of a failed response or thrown error into milliseconds so the engine can back off precisely. Return undefined when the error is not a throttle (or carries no hint).

      The default reads the standard Retry-After header (RFC 9110 §10.2.3 — the header a 429 and a 503 carry), in both its delay-seconds and HTTP-date forms, from wherever the HTTP client put it. That is not a heuristic and not vendor-specific: there is one correct reading of it, and every HTTP connector benefits from having it read.

      This used to return undefined unconditionally, and no connector in this repo overrode it — so the engine's limiter never learned a delay any vendor had actually stated, and discovery's throttle check (which asked this and only this) concluded "not a throttle" for every 429 MJ ever received.

      Override when the vendor signals its delay somewhere non-standard — in the response body, or in prose (PheedLoop's "Expected available in N second"). Deliberately not parsed here: guessing a duration out of message text risks inventing one, and a wrong Retry-After is worse than none, since it freezes the token bucket for a made-up interval.

      Parameters

      • error: unknown

      Returns number

    • Returns the ActionGeneratorConfig for this connector, combining the integration name, category, icon, and objects into a ready-to-use configuration for ActionMetadataGenerator.Generate().

      Override in subclasses to customize the config (e.g., icon, category). Returns null by default if GetIntegrationObjects() returns empty.

      Returns ActionGeneratorConfig

    • Returns suggested default field mappings for an external object to MJ entity. Override in subclasses to provide intelligent defaults.

      Parameters

      • _objectName: string

        Name of the external object

      • _entityName: string

        Name of the target MJ entity

      Returns DefaultFieldMapping[]

      Array of default field mappings (empty by default)

    • Returns the integration objects and their fields that this connector supports, for use by the ActionMetadataGenerator. This is static metadata that does NOT require a live connection — it describes the connector's known object model.

      Override in subclasses to provide connector-specific objects/fields. Returns an empty array by default (no action generation available).

      Returns IntegrationObjectInfo[]

    • Type-driven post-processing hook (plan.md §10): a connector may normalize/enforce a record's values to the resolved column formats AFTER transform/normalize and BEFORE write. Default returns the record unchanged. (Named for this system — NOT MCP, not take.) The engine ALSO applies target-type constraint enforcement; this is the connector-side complement.

      Parameters

      Returns ExternalRecord

    • Name of a stable, monotonic ordering key (PK/identity) usable for KEYSET/seek resume on watermark-less objects (plan.md §7 — resume from last-seen key, robust to mid-stream insert/delete). null → keyset resume unavailable for this object.

      Parameters

      • _objectName: string

      Returns string

    • Upserts a record — a single idempotent create-or-update keyed by a unique business property (e.g. email), eliminating the search-then-create race window. Override in subclasses whose external system exposes a keyed upsert primitive. Check SupportsUpsert before calling.

      Parameters

      Returns Promise<CRUDResult>