Documentation

Salesforce Connector

For a general introduction to the connector, please refer to RheinInsights Salesforce Enterprise Search and RAG Connector.

Salesforce Configuration

Crawl User

The connector uses OAuth and a (technical) crawl user which has the following permissions:

  1. Read access to all relevant content, which should be indexed

  2. Permission to the respective restrictions and permissions

  3. Read access to all users, groups and roles and their membership relationships

App Registration

Create an SSL Key Pair

openssl req -x509 -newkey rsa:2048 -nodes -days 730 \
  -keyout server.key -out server.crt -subj "/CN=rheininsights-salesforce-connector"

Please note that

  • server.crt is the public certificate and will be uploaded to Salesforce.

  • server.key stays with you;

  • leave it without a passphrase

  • make a note in your operations manual on how long the certificate is valid (730 days in this example)

Create the app in Salesforce (Setup)

  1. Click on the gearwheel and go to setup

  2. Search for External Client App Manager and open it

  3. Click on New external client app

    1. Name: RheinInsights RAG Connector

    2. Contact mail: no-reply@yourorganization.com

    3. API Name: Connector

    4. Distribution: Local

    5. Enable OAuth.

      1. Enter any callback URL, for example http://localhost:1717/OauthRedirect

      2. Add the following OAuth scopes:

        1. Manage user data via APIs (api) and

        2. Perform requests at any time (refresh_token, offline_access).

    6. Under Flow Enablement

      1. check the box at Enable JWT Bearer Flow

      2. Upload the public certificate “server.crt” from above

    7. Save. In the app's Policies, set Permitted Users to Admin approved users are pre-authorized, then add the integration user's profile or a permission set assigned to that user.

    8. Copy the Consumer Key from the OAuth settings.

  4. Add the App Policy

    1. In Setup, search for permission sets

    2. Click on new.

      1. Name it, for example, "RheinInsights RAG Connector"

      2. Leave the license empty

      3. Click save.

    3. In the new permission set, open assigned connected apps or External Client App Access.

    4. Click on edit

      1. Add connector

      2. Click manage assignments

      3. Add assignments and pick your user

  5.  Link the permission set to the app

    1. Go back to External Client App Manager → RheinInsights RAG Connector → Policies tab

      Click edit.

    2. In app policies with select permission sets:

      1. Move RheinInsights RAG Connector to the right under permission sets

      2. Click save.

  6. Fetch the consumer key:

    1. Also in External Client App Manager → RheinInsights RAG Connector

    2. Open Settings

    3. Expand OAuth settings and click consumer key and secret.

    4. Salesforce sends a verification code to your e-mail address. Enter it.

    5. Make a copy the consumer key, i.e. the one which starts with 3MVG.

      Consumer secret can be ignored.

Screenshot of the external client app registration

Content Source Configuration

The content source configuration of the connector comprises the following mandatory configuration fields.

Configuration Dialog

  1. Login URL. This is the login url for your salesforce instance.

  2. Integration user. This is the technical crawl user who is allowed to access the app which we configured above.

  3. Connected app consumer key. This is the consumer key as was generated in the last steps above.

  4. Private key. This is the private key which was generated above. It must fit the public key.

  5. API version. This is the Salesforce API version used by the connector.

  6. User identity field. This is the leading id field for the users. This is needed for security trimming reasons.

  7. Included Salesforce objects. This is a non-empty list of objects which is needed for crawling purposes. If you need more or less object types, then you can add these here.

  8. Excluded attachments from crawling: here you can add file extensions to filter documents which should not be sent to the search engine.

  9. Fetch binary content. When enabled, the binary body of Documents, Attachments and ContentVersions is fetched and indexed. When disabled, only metadata is indexed for these.

  10. Index Chatter feed and case comments. When enabled, comments on FeedItems and Cases are fetched and included in the indexed content.

  11. Add Chatter topics as metadata. When enabled, Chatter topics assigned to a record are added as metadata.

  12. Add tags as metadata. When enabled, tags assigned to a record are added as metadata.

  13. SOQL query page size. Row limit or page size for the SOQL result sets.

  14. Rate Limit. This will define a rate limiting for the connector, i.e., limit the number of API requests per second (across all threads).

  15. Page size for requests. Defines how many pages will be fetched per API request. Default is 100.

  16. Response timeout (ms). Defines how long the connector until an API call is aborted and the operation be marked as failed.

  17. Connection timeout (ms). Defines how long the connector waits for a connection for an API call.

  18. Socket timeout (ms). Defines how long the connector waits for receiving all data from an API call.

  19. The general settings are described at General Crawl Settings and you can leave these with its default values.

After entering the configuration parameters, click on validate. This validates the content crawl configuration directly against the content source. If there are issues when connecting, the validator will indicate these on the page. Otherwise, you can save the configuration and continue with Content Transformation configuration.

Recommended Crawl Schedules

Salesforce offers a partially complete change log. Only deletions and some types of modifications are not delivered in there. This means, we recommend to configure

  • incremental crawls every 2-4 hours

  • full scans to run every 24 hours,

  • and full scan principal crawls can run twice a day.

For more information see Crawl Scheduling .