A powerful, easily deployable network traffic analysis tool suite for network security monitoring
Malcolm’s runtime settings are stored (with a few exceptions) as environment variables in configuration files ending with a .env suffix in the ./config directory. The ./scripts/configure script can help users configure and tune these settings. For an in-depth treatment of the configuration script, see the Configuration section in End-to-end Malcolm and Hedgehog Linux ISO Installation.
Although the configuration script automates many of the following configuration and tuning parameters, some environment variables of particular interest are listed here for reference.
arkime.env and arkime-secret.env - settings for Arkime
ARKIME_DEBUG_LEVEL - debug flag for Arkime’s config.ini (see Arkime settings) (default 0)ARKIME_PCAP_LIBPCAP - if set to true, reverts to older libpcap mode for PCAP file processing rather than the faster scheme processing (default false)ARKIME_PASSWORD_SECRET - the password hash secret for the Arkime viewer cluster (see passwordSecret in Arkime INI Settings) used to secure the connection used when Arkime viewer retrieves a PCAP payload for display in its user interfaceARKIME_ROTATE_INDEX - how often (based on network traffic timestamp) to create a new index in OpenSearchARKIME_QUERY_ALL_INDICES - whether or not Arkime should query all indices instead of trying to calculate which ones pertain to the search time frame (default false)ARKIME_SPI_DATA_MAX_INDICES - the maximum number of indices for querying SPI data, or set to -1 to disable any max. The Arkime documentation warns “OpenSearch/Elasticsearch MAY blow up if we … search too many indices.” (default 7)MANAGE_PCAP_FILES and ARKIME_FREESPACEG - these variables deal with PCAP deletion by Arkime, see Managing disk usage belowMAXMIND_GEOIP_DB_ACCOUNT_ID and MAXMIND_GEOIP_DB_LICENSE_KEY - Malcolm uses MaxMind’s free GeoLite2 databases for GeoIP lookups. As of December 30, 2019, these databases are no longer available for download via a public URL. Instead, they must be downloaded using a MaxMind license key (available without charge from MaxMind). The MaxMind acount ID and license key can be specified here for GeoIP database downloads during build- and run-time.MAXMIND_GEOIP_DB_ALTERNATE_DOWNLOAD_URL - As an alternative to (or fallback for) MAXMIND_GEOIP_DB_LICENSE_KEY, a URL prefix may be specified in this variable (e.g., https://example.org/foo/bar) which will be used as a fallback. This URL should serve up .tar.gz files in the same format as those provided by the official source (see the example here).INDEX_MANAGEMENT_ENABLED - if set to true, Malcolm’s instance of Arkime will use these features when indexing data; note that this only takes effect when initializing Malcolm from an empty stateINDEX_MANAGEMENT_OPTIMIZATION_PERIOD - the period in hours or days that Arkime will keep records in the hot state (default 30d)INDEX_MANAGEMENT_RETENTION_TIME - the period in hours or days that Arkime will keep records before deleting them (default 90d)INDEX_MANAGEMENT_OLDER_SESSION_REPLICAS - the number of replicas for older sessions indices (default 0)INDEX_MANAGEMENT_HISTORY_RETENTION_WEEKS - the retention time period (weeks) for Arkime history data (default 13)INDEX_MANAGEMENT_SEGMENTS - the number of segments Arkime will use to optimize sessions (default 1)INDEX_MANAGEMENT_HOT_WARM_ENABLED - whether or not Arkime should use a hot/warm design (storing non-session data in a warm index); setting up hot/warm index policies also requires configuration on the local nodes in accordance with the Arkime documentationARKIME_WISE_CONFIG_PIN_CODE - the WISE service requires a configuration pin if the wise UI is write-enabled via ARKIME_EXPOSE_WISE_GUI (see below). This value will be required to save any WISE configuration changes. The default value is WISE2019.ARKIME_WISE_SERVICE_URL - to leverage WISE, Arkime capture needs to be provided a wiseURL value. The value of this environment variable is copied into the wiseURL value in arkime-live containers.arkime-live.env - settings for live traffic capture with Arkime
ARKIME_LIVE_CAPTURE - whether or not Arkime should monitor live traffic on a local interface (PCAP_IFACE in pcap-capture.env specifies the interface)ARKIME_LIVE_NODE_HOST - the node host (e.g., the IP address of the machine running Malcolm) to associate with network traffic metadata when ARKIME_LIVE_CAPTURE is true (optional, defaults to PCAP_NODE_NAME if unspecified)WISE - indicates if the WISE service is on or off. This environment variable defaults to off.ARKIME_COMPRESSION_TYPE, ARKIME_COMPRESSION_LEVEL, ARKIME_DB_BULK_SIZE, ARKIME_MAGIC_MODE, ARKIME_MAX_PACKETS_IN_QUEUE, ARKIME_PACKET_THREADS, ARKIME_PCAP_WRITE_METHOD, ARKIME_PCAP_WRITE_SIZE, ARKIME_PCAP_READ_METHOD, ARKIME_TPACKETV3_NUM_THREADS, and ARKIME_TPACKETV3_BLOCK_SIZE, all of which relate to managing Arkime’s performance and resource utilization during live capture.arkime-offline.env - settings for analyzing uploaded PCAP files with Arkime
ARKIME_AUTO_ANALYZE_PCAP_FILES - if set to true, all PCAP files imported into Malcolm will automatically be analyzed by Arkime, and the resulting logs will also be imported (default false)ARKIME_AUTO_ANALYZE_PCAP_THREADS - the number of threads available to Arkime for analyzing PCAP files (default 1)ARKIME_ROTATED_PCAP - whether or not Arkime should analyze captured PCAP files captured by netsniff-ng/tcpdump (see PCAP_ENABLE_NETSNIFF and PCAP_ENABLE_TCPDUMP below). If ARKIME_LIVE_CAPTURE is true, this should be false; otherwise Arkime will see duplicate traffic.ARKIME_EXPOSE_WISE_GUI - if set to true the WISE interface will be available at: https://<MALCOLM-IP>/wise. This defaults to true. Note that the WISE configuration GUI is only run in the arkime container on a Malcolm instance running the malcolm run profile.ARKIME_ALLOW_WISE_GUI_CONFIG - if set to true the WISE interface can be used to configure the WISE service, i.e., the WISE configuration GUI is editable rather than read-only. This only applies if ARKIME_EXPOSE_WISE_GUI is set to true. The default value is true.auth-common.env - authentication-related settings
NGINX_AUTH_MODE - valid values are basic (or true for legacy compatibility), use TLS-encrypted HTTP basic authentication (default); ldap (or false for legacy compatibility), use Lightweight Directory Access Protocol (LDAP) authentication; keycloak to use authentication managed by Malcolm’s embedded Keycloak instance; keycloak_remote to use authentication managed by a remote Keycloak instance; no_authentication to disable authenticationNGINX_REQUIRE_GROUP and NGINX_REQUIRE_ROLE - When using Keycloak authentication, setting these values will require authenticated users to belong to groups and assigned roles, respectively. Multiple values may be specified with a comma-separated list. Note that these requirements are cumulative: users must match all of the items specified. An empty value (default) means no group/role restriction is applied. LDAP authentication can also require group membership, but that is specified in nginx_ldap.conf by setting require group rather than in auth-common.env.ROLE… - variables used to manage role-based access controlauth.env - stores the Malcolm administrator’s username and password hash for its nginx reverse proxybeats-common.env - settings for interactions between Logstash and Filebeat
BEATS_SSL - if set to true, Logstash will use require encrypted communications for any external Beats-based forwarders from which it will accept logs (default true)LOGSTASH_HOST - the host and port at which Beats-based forwarders will connect to Logstash (default logstash:5044); see MALCOLM_PROFILE belowdashboards.env and dashboards-helper.env - settings for the containers that configure and maintain OpenSearch and OpenSearch Dashboards
DASHBOARDS_URL - used primarily when OPENSEARCH_PRIMARY is set to elasticsearch-remote (see OpenSearch and Elasticsearch instances), this variable stores the URL for the Kibana instance into which Malcolm’s dashboard’s and index templates will be importedDASHBOARDS_PREFIX - a string to prepend to the titles of Malcolm’s prebuilt dashboards prior upon import during Malcolm’s initialization (default is an empty string)DASHBOARDS_DARKMODE - if set to true, OpenSearch Dashboards will be set to dark mode upon initialization (default true)DASHBOARDS_DATA_SOURCES_ENABLED - Enable/disable Dashboards’ multiple data sources feature, letting one Dashboards instance query and visualize data from several separate clusters at once. Users may turn this on when running a central Dashboards instance that needs to reach into more than one Malcolm OpenSearch backend; leave it false (the default) for a standard single-cluster deployment. This setting is experimental.DASHBOARDS_TIMEPICKER_FROM and DASHBOARDS_TIMEPICKER_TO - sets the “from” and “to” values, respectively, for OpenSearch Dashboard’s timepicker:timeDefaults setting (default now-24h and now, meaning “last 24 hours”)OPENSEARCH_INDEX_SIZE_PRUNE_LIMIT - the maximum cumulative size of OpenSearch indices are allowed to consume before the oldest indices are deleted, see Managing disk usage belowOPENSEARCH_INDEX_SIZE_PRUNE_NAME_SORT - whether to determine the “oldest” indices for storage-based index pruning by creation date/time (false, default) or index name (true)OPENSEARCH_DEFAULT_DASHBOARD - the dashboard to open by default, specified by its ID (default is Malcolm’s Overview dashboard)filebeat.env - settings specific to Filebeat, particularly for how Filebeat watches for new log files to parse and how it receives and stores third-Party logs
LOG_CLEANUP_MINUTES and ZIP_CLEANUP_MINUTES - these variables deal cleaning up already-processed log files, see Managing disk usage belowFILEBEAT_CLEANUP_VERBOSITY - verbosity flag for the log/zip file cleanup process (e.g., -v, -vv, -vvv, etc.)FILEBEAT_SCAN_FREQUENCY, FILEBEAT_CLEAN_INACTIVE, FILEBEAT_IGNORE_OLDER, FILEBEAT_CLOSE_INACTIVE, FILEBEAT_CLOSE_INACTIVE_LIVE, FILEBEAT_CLOSE_RENAMED, FILEBEAT_CLOSE_REMOVED, FILEBEAT_CLOSE_EOF, FILEBEAT_CLEAN_REMOVED - Filebeat parameters used for monitoring log files containing network traffic metadata (see Filebeat log input settings)FILEBEAT_SCANNER_FINGERPRINT_OFFSET and FILEBEAT_SCANNER_FINGERPRINT_LENGTH - the file offset and fingerprint length, respectively, for .prospector.scanner.fingerprintFILEBEAT_TCP_LISTEN - whether or not to expose a filebeat TCP input listener (see Filebeat TCP input)FILEBEAT_TCP_LOG_FORMAT - log format expected for events sent to the filebeat TCP input listener (json or raw)FILEBEAT_TCP_PARSE_SOURCE_FIELD - source field name to parse (when FILEBEAT_TCP_LOG_FORMAT is json) for events sent to the filebeat TCP input listenerFILEBEAT_TCP_PARSE_TARGET_FIELD - target field name to store decoded JSON fields (when FILEBEAT_TCP_LOG_FORMAT is json) for events sent to the filebeat TCP input listenerFILEBEAT_TCP_PARSE_DROP_FIELD - name of a field to drop (if it exists) in events sent to the filebeat TCP input listenerFILEBEAT_TCP_TAG - tag to append to events sent to the filebeat TCP input listenerFILEBEAT_PREPARE_PROCESS_COUNT - number of processes dedicated to preparing files for ingestion into filebeatFILEBEAT_SYSLOG_TCP_LISTEN and FILEBEAT_SYSLOG_UDP_LISTEN - if set to true, Malcolm will accept syslog messages over TCP and/or UDP, respectivelyFILEBEAT_SYSLOG_TCP_PORT and FILEBEAT_SYSLOG_UDP_PORT - the port on which Malcolm will accept syslog messages over TCP and/or UDP, respectively
FILEBEAT_SYSLOG_TCP_FORMAT and FILEBEAT_SYSLOG_UDP_FORMAT - one of auto, rfc3164, or rfc5424, to specify the allowed format for syslog messages over TCP and/or UDP, respectively (default auto)FILEBEAT_SYSLOG_TCP_MAX_MESSAGE_SIZE and FILEBEAT_SYSLOG_UDP_MAX_MESSAGE_SIZE - defines the maximum message size of the message received over TCP and/or UDP, respectively (default: 10KiB for UDP, 20MiB for TCP)FILEBEAT_SYSLOG_TCP_MAX_CONNECTIONS - specifies the maximum current number of TCP connections for syslog messagesFILEBEAT_SYSLOG_TCP_SSL - if set to true, syslog messages over TCP will require the use of TLS. When ./scripts/auth_setup is run, self-signed certificates are generated which may be used by remote log forwarders. Located in the filebeat/certs/ directory, the certificate authority and client certificate and key files should be copied to the host on which the forwarder is running and used when defining its settings for connecting to Malcolm.filescan.env, filescan-secret.env, and pipeline.env - Settings related to scanning of automatically-extracted files observed in traffic using Strelka (see also zeek.env below)
CLAMD_… - variables used to managed the ClamAV server to which Strelka can submit files for scanningFILESCAN_HTTP_SERVER_ENABLE - if set to true, the directory containing Zeek-extracted files will be served over HTTP at ./extracted-files/ (e.g., https://localhost/extracted-files/ if connecting locally)FILESCAN_HTTP_SERVER_KEY - specifies the password for the ZIP archive if FILESCAN_HTTP_SERVER_ZIP is true; otherwise, this specifies the decryption password for encrypted Zeek-extracted files in an openssl enc-compatible format (e.g., openssl enc -aes-256-cbc -d -in example.exe.encrypted -out example.exe)FILESCAN_HTTP_SERVER_ZIP - if to true, the Zeek-extracted files will be archived in a ZIP file upon downloadFILESCAN_HTTP_SERVER_MAGIC - whether or not to use libmagic to show MIME types for Zeek-extracted files served over HTTPFILESCAN_PRESERVATION - determines behavior for preservation of Zeek-extracted filesFILESCAN_PRUNE_INTERVAL_SECONDS - the interval between checking the prune conditions, in seconds (default 300)FILESCAN_PRUNE_THRESHOLD_MAX_SIZE - specifies the maximum size, specified either in gigabytes or as a human-readable data size (e.g., 250G), that the ./zeek-logs/extract_files/ directory is allowed to contain before the prune condition triggersFILESCAN_PRUNE_THRESHOLD_TOTAL_DISK_USAGE_PERCENT - specifies a maximum fill percentage for the file system containing the ./zeek-logs/extract_files/; in other words, if the disk is more than this percentage utilized, the prune condition triggersPIPELINE_DISABLED - if set to true, file scanning with Strelka will be disabledRULES_UPDATE_ENABLED - if set to true, file scanner engines (e.g., ClamAV, YARA, etc.) will periodically update their rule definitions (default false)STRELKA_BACKEND_PROCS - specifies the number of Strelka backend instances performing file scanning concurrently (default 1)STRELKA_SCANNERS - comma-separated list which may contain scanner names and/or the string default for the built-in list of scanners (see MALCOLM_STRELKA_SCANNERS_DEFAULT in malcolm_constants.py). Values may be combined (e.g., default,ScanStrings).YARA_CUSTOM_RULES_ONLY - if set to true, Malcolm will bypass its default YARA rulesets and use only user-defined rules in ./yara/rules (default false)keycloak.env - settings specific to Keycloak
NGINX_AUTH_MODE above):
KEYCLOAK_AUTH_REALM - specifies the name of the Keycloak realm (default master)KEYCLOAK_AUTH_REDIRECT_URI - specifies the relative path which is the Malcolm URI to which Keycloak will redirect users after a successful authentication (default /index.html, which will redirect users to the Malcolm landing page)KEYCLOAK_AUTH_URL - specifies the Keycloak endpoint URL, or the URL to which Malcolm should direct authentication requests for Keycloak. If a remote Keycloak instance is being used, this would be the URL for that instance (e.g., https://keycloak.example.com ). If Malcolm is using its embedded Keycloak instance, this host portion of the URL should be the hostname or IP address at which Malcolm is available, followed by /keycloak (or whatever the value of KC_HTTP_RELATIVE_PATH has been set to; see below) (e.g., https://malcolm.internal.lan/keycloak or https://192.168.100.10/keycloak ).KEYCLOAK_CLIENT_ID and KEYCLOAK_CLIENT_SECRET - identify the Keycloak client Malcolm will use and the secret associated with that clientKEYCLOAK_SSL_VERIFY - if set to true, requires SSL certificate verification when connecting to a remote Keycloak endpoint (used when NGINX_AUTH_MODE is keycloak_remote)`KC_CACHE - defines the cache mechanism for high-availability (default local as Malcolm’s embedded Keycloak instance is single-node)KC_HEALTH_ENABLED - if set to true, enables the health check endpoint used internally by the container health scriptKC_HOSTNAME - address at which the Keycloak server is exposed (defaults to blank, as KC_HOSTNAME_STRICT below defaults to false)KC_HOSTNAME_STRICT - if set to true, disables dynamically resolving the hostname from request headersKC_HTTP_ENABLED - enables the HTTP listener (default true as Malcolm is proxying the embedded Keycloak instance behind nginx)KC_HTTP_RELATIVE_PATH - specifies the Malcolm path under which Keycloak serves resources (should not be changed from its default value of /keycloak)KC_METRICS_ENABLED - specifies if the server should expose metrics (default false)KC_PROXY_HEADERS - the proxy headers that should be accepted by Keycloak (should not be changed from its default value of xforwarded)KC_BOOTSTRAP_ADMIN_USERNAME and KC_BOOTSTRAP_ADMIN_PASSWORD - values for bootstrapping the temporary Keycloak admin service account (see Keycloak configuration)logstash.env - settings specific to Logstash
LOGSTASH_OUI_LOOKUP - if set to true, Logstash will map MAC addresses to vendors for all source and destination MAC addresses when analyzing Zeek logs (default true)LOGSTASH_REVERSE_DNS - if set to true, Logstash will perform a reverse DNS lookup for all external source and destination IP address values when analyzing Zeek logs (default false)LOGSTASH_SEVERITY_SCORING - if set to true, Logstash will perform severity scoring when analyzing Zeek logs (default true)LOGSTASH_NETBOX_ENRICHMENT_DATASETS - defines which types of logs will be enriched via NetBox: a comma-separated list which may contain provider.dataset pairs (e.g., zeek.dns), the string default for the built-in list of log types, ics (or ot) to enrich OT/ICS traffic, or all to enrich all logs. Values may be combined (e.g., default,ics).LOGSTASH_ZEEK_IGNORED_LOGS - a comma-separated list of Zeek log types that will be ignored (dropped) by LogstashLS_JAVA_OPTS - part of LogStash’s JVM settings, the -Xmx and -Xms values set the size of LogStash’s Java heap (we recommend somewhere between 1500m and 4g)pipeline.workers, pipeline.batch.size and pipeline.batch.delay - these settings are used to tune the performance and resource utilization of the logstash container; see Tuning and Profiling Logstash Performance, logstash.yml and Multiple PipelinesOTKB_ENRICHMENT - if set to true, Logstash will enrich OT/ICS network traffic metadata via the MITRE OT Knowledge Base (see also logstash-secret.env below)OTKB_ENRICHMENT_TTL - the time-to-live (in seconds) for the cached OTKB lookup data before it is refreshedOTKB_SSL_VERIFY - if set to false, disables SSL certificate verification when connecting to the OTKB APIOTKB_ENRICHMENT_VERBOSE - if set to true, includes the complete expanded OTKB records in enriched events; otherwise, only the recommended fields and fields used by Malcolm’s OTKB dashboard are includedOTKB_ENRICHMENT_DEBUG - if set to true, logs each OTKB API call for debuggingOTKB_ENRICHMENT_DEBUG_TIMINGS - if set to true, collects and logs timing statistics for OTKB API callslogstash-secret.env - secrets for enriching with the MITRE OT Knowledge Base (see also logstash.env above)
OTKB_URL - the base URL of the OTKB API used for enrichment lookups (e.g., https://otkb.example.org:31443/api/v1)OTKB_TOKEN - the API token used to authenticate to the OTKB API (sent as an Authorization: Token <token> header)lookup-common.env - settings for enrichment lookups, including those used for customizing event severity scoring
FREQ_LOOKUP - if set to true, domain names (from DNS queries and SSL server names) will be assigned entropy scores as calculated by freq (default false)FREQ_SEVERITY_THRESHOLD - when severity scoring is enabled, this variable indicates the entropy threshold for assigning severity to events with entropy scores calculated by freq; a lower value will only assign severity scores to fewer domain names with higher entropy (e.g., 2.0 for NQZHTFHRMYMTVBQJE.COM), while a higher value will assign severity scores to more domain names with lower entropy (e.g., 7.5 for naturallanguagedomain.example.org) (default 2.0)SENSITIVE_COUNTRY_CODES - when severity scoring is enabled, this variable defines a comma-separated list of sensitive countries (using ISO 3166-1 alpha-2 codes) (default 'AM,AZ,BY,CN,CU,DZ,GE,HK,IL,IN,IQ,IR,KG,KP,KZ,LY,MD,MO,PK,RU,SD,SS,SY,TJ,TM,TW,UA,UZ', taken from the U.S. Department of Energy Sensitive Country List)TOTAL_MEGABYTES_SEVERITY_THRESHOLD - when severity scoring is enabled, this variable indicates the size threshold (in megabytes) for assigning severity to large connections or file transfers (default 1000)netbox-common.env, netbox.env and netbox-secret.env - settings related to NetBox and Asset Interaction Analysis
NETBOX_MODE - determine whether Malcolm will start and manage a NetBox instance; valid values are local (use an embedded instance NetBox), remote (use a remote instance of NetBox), or disabled (the default)NETBOX_ENRICHMENT - if set to true, Logstash will enrich network traffic metadata via NetBox API callsNETBOX_ENRICHMENT_LOOKUP_SERVICE - whether or not services (i.e., destination IP/port) will be looked up during NetBox enrichment (default true)NETBOX_DEFAULT_SITE - specifies the default NetBox site name for use when enriching network traffic metadata via NetBox lookups if a specific site is not otherwise specified for the source of the data (default Malcolm)NETBOX_AUTO_POPULATE - if set to true, Logstash will populate the NetBox inventory based on observed network trafficNETBOX_AUTO_POPULATE_SUBNETS - a comma-separated list of private CIDR subnets to control NetBox IP autopopulation (see Subnets considered for autopopulation; default is an empty string, meaning all private IPv4 and IPv6 ranges are autopopulated)NETBOX_AUTO_CREATE_PREFIX - if set to true, Logstash will automatically create private subnet prefixes in the NetBox inventory based on observed network trafficNETBOX_DEFAULT_AUTOCREATE_MANUFACTURER - if set to true, new manufacturer entries will be created in the NetBox database when matching device manufacturers to OUIs (default true)NETBOX_DEFAULT_FUZZY_THRESHOLD - fuzzy-matching threshold for matching device manufacturers to OUIs (default 0.95)NETBOX_CACHE_SIZE and NETBOX_CACHE_TTL - caching parameters for NetBox’s Logstash lookupsNETBOX_MODE is set to remote; otherwise, they should be blank:
NETBOX_URL - the URL of the remote NetBox instance (e.g., https://netbox.example.org or https://example.com/netbox)NETBOX_TOKEN - the API token for the remote NetBox instance (40 hexadecimal characters)nginx.env - settings specific to Malcolm’s nginx reverse proxy
NGINX_LOG_ACCESS_AND_ERRORS - if set to true, all access to Malcolm via its web interfaces will be logged to OpenSearch (default false)NGINX_ERROR_LOG_LEVEL - sets the logging level for nginx’s error log (e.g., debug, info, notice, warn, error, crit, alert, emerg; default error)NGINX_SSL - if set to true, require HTTPS connections to Malcolm’s nginx-proxy container (default); if set to false, use unencrypted HTTP connections (using unsecured HTTP connections is NOT recommended unless you are running Malcolm behind another reverse proxy such as Traefik, Caddy, etc.) Also note: in some circumstances disabling SSL in NGINX while leaving SSL enabled in Arkime can result in a “Missing token” Arkime error. This is due to Arkime’s Cross-Site Request Forgery mitigation cookie being passed to the browser with the “secure” flag enabled.NGINX_X_FORWARDED_PROTO_OVERRIDE - overrides the scheme (http/https) used when constructing the OIDC redirect URI. Set this when nginx is behind a TLS-terminating reverse proxy or ingress controller where the internal connection is HTTP (e.g., NGINX_SSL=false) but the client-facing connection is HTTPS. If unset, the scheme is inferred normally.NGINX_CSP_FORM_ACTION_EXTRA - space-separated list of additional CSP source expressions permitted as form submission targets, typically identity provider origins (e.g., https://idp.example.org). Set this when an authentication flow, such as SAML HTTP-POST binding, must submit a form to an external identity provider. Do not separate entries with commas or surround individual origins with quotes. See related issue.NGINX_AUTH_MODE=ldap) support for LDAP, LDAPS, or LDAP+StartTLS connections:
NGINX_LDAP_TLS_STUNNEL - for StartTLS, set this to true to issue the StartTLS command and use stunnel to tunnel the connectionNGINX_LDAP_TLS_STUNNEL_CHECK_HOST and NGINX_LDAP_TLS_STUNNEL_CHECK_IP - stunnel will require and verify certificates for StartTLS when one or more trusted CA certificate files are placed in the ./nginx/ca-trust directory. For additional security, hostname or IP address checking of the associated CA certificate(s) can be enabled by providing these values.NGINX_LDAP_TLS_STUNNEL_VERIFY_LEVEL - sets stunnel’s certificate verification level for the StartTLS connection (0 = no verification, 1 = verify peer certificate if present, 2 = verify peer certificate, 3 = verify with locally-installed certificate; default 2)NGINX_RESOLVER_OVERRIDE - if set, overrides automatic detection of the resolver address used (default is unset)NGINX_RESOLVER_IPV4 - if false, sets the ipv4=off parameter in the resolver directive (default is true)NGINX_RESOLVER_IPV6 - if false, sets the ipv6=off parameter in the resolver directive; it is recommended to set this to false if your network does not support IPv6 (default is true)opensearch.env - settings specific to OpenSearch
OPENSEARCH_JAVA_OPTS - one of OpenSearch’s most important settings, the -Xmx and -Xms values set the size of OpenSearch’s Java heap (we recommend setting this value to half of system RAM, up to 32 gigabytes)OPENSEARCH_PRIMARY - one of opensearch-local, opensearch-remote, or elasticsearch-remote, to determine the OpenSearch or Elasticsearch instance Malcolm will use (default opensearch-local)OPENSEARCH_URL - when using Malcolm’s internal OpenSearch instance (i.e., OPENSEARCH_PRIMARY is opensearch-local) this should be https://opensearch:9200; otherwise, this value specifies the primary remote instance URL in the format protocol://host:port (default https://opensearch:9200)OPENSEARCH_SSL_CERTIFICATE_VERIFICATION - if set to true, connections to the primary remote OpenSearch instance will require full TLS certificate validation (this may fail if using self-signed certificates) (default false)OPENSEARCH_SECONDARY - one of opensearch-local, opensearch-remote, elasticsearch-remote, or blank (unset) to indicate that Malcolm should forward logs to a secondary remote OpenSearch instance in addition to the primary OpenSearch instance (default is unset)OPENSEARCH_SECONDARY_URL - when forwarding to a secondary remote OpenSearch instance (i.e., OPENSEARCH_SECONDARY is set) this value specifies the secondary remote instance URL in the format protocol://host:portOPENSEARCH_SECONDARY_SSL_CERTIFICATE_VERIFICATION - if set to true, connections to the secondary remote OpenSearch instance will require full TLS certificate validation (this may fail if using self-signed certificates) (default false)MALCOLM_NETWORK_INDEX_PATTERN - Index pattern for network traffic logs written via Logstash (default is arkime_sessions3-*)MALCOLM_NETWORK_INDEX_ALIAS - specifies the value under template.aliases in malcolm_template, or blank to not specify oneMALCOLM_NETWORK_INDEX_DEFAULT_PIPELINE - specifies the value of template.settings.index.default_pipeline in malcolm_template, or blank to not specify oneMALCOLM_NETWORK_INDEX_LIFECYCLE_NAME - specifies the value of template.settings.index.lifecycle.name in malcolm_template, or blank to not specify one (Elasticsearch only)MALCOLM_NETWORK_INDEX_LIFECYCLE_ROLLOVER_ALIAS - specifies the value of template.settings.index.lifecycle.rollover_alias in malcolm_template, or blank to not specify one (Elasticsearch only)MALCOLM_NETWORK_INDEX_TIME_FIELD - Default time field to use for network traffic logs in Logstash and Dashboards (default is firstPacket)MALCOLM_NETWORK_INDEX_SUFFIX - Suffix used to create index to which network traffic logs are written
strftime strings in %{}) (e.g., hourly: %{%y%m%dh%H}, twice daily: %{%P%y%m%d}, daily (default): %{%y%m%d}, weekly: %{%yw%U}, monthly: %{%ym%m}{{ }} (e.g., {{event.provider}}%{%y%m%d})MALCOLM_OTHER_INDEX_PATTERN - Index pattern for other logs written via Logstash (default is malcolm_beats_*)MALCOLM_OTHER_INDEX_ALIAS - specifies the value under template.aliases in malcolm_beats_template, or blank to not specify oneMALCOLM_OTHER_INDEX_DEFAULT_PIPELINE - specifies the value of template.settings.index.default_pipeline in malcolm_beats_template, or blank to not specify oneMALCOLM_OTHER_INDEX_LIFECYCLE_NAME - specifies the value of template.settings.index.lifecycle.name in malcolm_beats_template, or blank to not specify one (Elasticsearch only)MALCOLM_OTHER_INDEX_LIFECYCLE_ROLLOVER_ALIAS - specifies the value of template.settings.index.lifecycle.rollover_alias in malcolm_beats_template, or blank to not specify one (Elasticsearch only)MALCOLM_OTHER_INDEX_TIME_FIELD - Default time field to use for other logs in Logstash and Dashboards (default is @timestamp)MALCOLM_OTHER_INDEX_SUFFIX - Suffix used to create index to which other logs are written (with the same rules as MALCOLM_NETWORK_INDEX_SUFFIX above) (default is %{%y%m%d})ARKIME_NETWORK_INDEX_PATTERN and ARKIME_NETWORK_INDEX_TIME_FIELD - the index pattern and default time field, respectively, used specifically by Arkime (will probably match MALCOLM_NETWORK_INDEX_PATTERN/MALCOLM_NETWORK_INDEX_TIME_FIELD above)MALCOLM_INDEX_MAX_RESULT_WINDOW - if specified, overrides index.max_result_window, which defines the maximum value of from + size for searches of the network data index. CAUTION: increasing this value beyond its baked-in default (10000) may negatively impact performance.ARKIME_INIT_SHARDS, ARKIME_INIT_REPLICAS, ARKIME_INIT_REFRESH_SEC, ARKIME_INIT_SHARDS_PER_NODE - these variables, if specified, are used for Arkime’s <init opts> for db.pl init and upgrade operations and for the corresponding index pattern template created during setupCLUSTER_MAX_SHARDS_PER_NODE - sets OpenSearch’s persistent cluster setting cluster.max_shards_per_node, the maximum number of shards allowed per node (default 2500)MAX_LOCKED_MEMORY - the memlock limit for the OpenSearch container, used in conjunction with bootstrap.memory_lock to let OpenSearch lock its heap into memory and avoid it being swapped out (default unlimited)pcap-capture.env - settings specific to capturing traffic for live traffic analysis
PCAP_ENABLE_NETSNIFF - if set to true, Malcolm will capture network traffic on the local network interface(s) indicated in PCAP_IFACE using netsniff-ngPCAP_ENABLE_TCPDUMP - if set to true, Malcolm will capture network traffic on the local network interface(s) indicated in PCAP_IFACE using tcpdump; there is no reason to enable both PCAP_ENABLE_NETSNIFF and PCAP_ENABLE_TCPDUMPPCAP_FILTER - specifies a tcpdump-style filter expression for local packet capture; leave blank to capture all trafficPCAP_IFACE - used to specify the network interface(s) for local packet capture if PCAP_ENABLE_NETSNIFF, PCAP_ENABLE_TCPDUMP, ZEEK_LIVE_CAPTURE or SURICATA_LIVE_CAPTURE are enabled; for multiple interfaces, separate the interface names with a comma (e.g., 'enp0s25' or 'enp10s0,enp11s0')PCAP_IFACE_TWEAK - if set to true, Malcolm will use ethtool to disable NIC hardware offloading features and adjust ring buffer sizes for capture interface(s); this should be true if the interface(s) are being used for capture only, false if they are being used for management/communicationPCAP_ROTATE_MEGABYTES - used to specify how large a locally captured PCAP file can become (in megabytes) before it is closed for processing and a new PCAP file createdPCAP_ROTATE_MINUTES - used to specify a time interval (in minutes) after which a locally-captured PCAP file will be closed for processing and a new PCAP file createdPCAP_IFACE_STATS_CRON_EXPRESSION - Specifies a cron expression (using cronexpr-compatible syntax) indicating the refresh interval for collecting kernel-level statistics for network interfaces. An empty value for this variable means these statistics will not be generated.postgres.env - Internal settings related to the PostgreSQL relational databaseprocess.env - settings for how the processes running inside Malcolm containers are executed
PUID and PGID - Docker runs all its containers as the privileged root user by default. For better security, Malcolm immediately drops to non-privileged user accounts for executing internal processes wherever possible. The PUID (process user ID) and PGID (process group ID) environment variables allow Malcolm to map internal non-privileged user accounts to a corresponding user account on the host. Note a few (including the logstash and netbox containers) may take a few extra minutes during startup if PUID and PGID are set to values other than the default 1000. This is expected and should not affect operation after the initial startup.MALCOLM_PROFILE - Specifies the profile which determines the Malcolm containers to run (malcolm to run all containers, hedgehog to run only capture-related containers)MALCOLM_CONTAINER_RUNTIME - specifies the container runtime engine (e.g., docker, podman)TINI_VERBOSITY - for debugging container init via tinissl.env - TLS-related settings used by many containers
PUSER_CA_TRUST - when possible, docker containers will automatically add trusted CA certificate files found in the ./nginx/ca-trust directory (which is bind mounted to /ca-trust)suricata.env, suricata-live.env and suricata-offline.env - settings for Suricata
SURICATA_AUTO_ANALYZE_PCAP_FILES - if set to true, all PCAP files imported into Malcolm will automatically be analyzed by Suricata, and the resulting logs will also be imported (default false)SURICATA_AUTO_ANALYZE_PCAP_PROCESSES - the number of processes available to Malcolm for processing PCAP files with Suricata (default 1)SURICATA_AUTO_ANALYZE_PCAP_THREADS - the number of threads to use per Suricata process (default 0, meaning Suricata will use its default behavior)SURICATA_CUSTOM_RULES_ONLY - if set to true, Malcolm will bypass the default Suricata ruleset and use only user-defined rules (./suricata/rules/*.rules).SURICATA_UPDATE_RULES - if set to true, Suricata signatures will periodically be updated (default false)SURICATA_UPDATE_SOURCES - whether or not to refresh the suricata-update source index before updating rules (default true)SURICATA_UPDATE_DEBUG - if set to true, runs suricata-update with verbose output instead of quiet mode (default false)SURICATA_UPDATE_CRON_EXPRESSION - a five-field cron expression controlling when the Suricata rule update job runsSURICATA_UPDATE_ENABLE_SOURCES - comma-separated suricata-update rules sources (from the Suricata Rule Index) to enable (default et/open). Optional source parameters may follow the source name separated by |, e.g. et/open,ptresearch/attackdetection,et/pro|secret-code=REDACTED. Named sources require a populated local source index, so leave SURICATA_UPDATE_SOURCES enabled for at least one successful source refresh before disabling index updates on a fresh volume. Source parameters are passed to suricata-update as argv values and can be visible in the process list while enable-source runs; avoid placing long-lived credentials here unless that exposure is acceptable.SURICATA_LIVE_CAPTURE - if set to true, Suricata will monitor live traffic on the local interface(s) defined by PCAP_FILTERSURICATA_RUNMODE - specifies the Suricata runmode for live capture (see Suricata runmodes)SURICATA_ROTATED_PCAP - if set to true, Suricata can analyze PCAP files captured by netsniff-ng or tcpdump (see PCAP_ENABLE_NETSNIFF and PCAP_ENABLE_TCPDUMP, as well as SURICATA_AUTO_ANALYZE_PCAP_FILES); if SURICATA_LIVE_CAPTURE is true, this should be false; otherwise Suricata will see duplicate trafficSURICATA_DISABLE_ICS_ALL - if set to true, this variable can be used to disable Malcolm’s built-in Suricata rules for Operational Technology/Industrial Control Systems (OT/ICS) vulnerabilities and exploitsSURICATA_DISABLE_SIDS - may be set to a comma-separated list of entries (e.g., rule sids) with which to populate Suricata’s disable.confSURICATA_STATS_ENABLED, SURICATA_STATS_EVE_ENABLED, SURICATA_STATS_INTERVAL, and SURICATA_STATS_DECODER_EVENTS - these variables control the generation of live traffic capture statistics for Suricata, which data is used to populate the Packet Capture Statistics dashboardupload-common.env - settings for dealing with PCAP files uploaded to Malcolm for analysis
AUTO_TAG - if set to true, Malcolm will automatically create Arkime sessions and Zeek logs with tags based on the filename, as described in Tagging (default true)EXTRA_TAGS - a comma-separated list of default tags for data generated by Malcolm (default is an empty string)PCAP_NODE_NAME - specifies the node name to associate with network traffic metadataPCAP_UPLOAD_MAX_FILE_GB - specifies the maximum uploadable file size in whole gigabytes (default 50)PCAP_PIPELINE_VERBOSITY - verbosity flag for pcap pipeline debugging (e.g., -v, -vv, -vvv, etc.)PCAP_PIPELINE_IGNORE_PREEXISTING - whether or not PCAP files extant in ./pcap/ will be ignored on startupPCAP_MONITOR_HOST - pcap-monitor, to match the name of the container providing the uploaded/captured PCAP file monitoring serviceSAFE_EXTRACT_MAX_ENTRIES, SAFE_EXTRACT_MAX_DEPTH, and SAFE_EXTRACT_MAX_BYTES - Malcolm supports uploading archive files containing Zeek logs and Microsoft Windows event log files (with a .evtx file extension). These SAFE_EXTRACT_… variables guard against “zip bombs.”valkey.env - Settings related to the Valkey in-memory database
VALKEY_MAXMEMORY - the maximum amount of memory the valkey service is allowed to use, e.g. 100mb (0 means no limit; default 0)VALKEY_MAXMEMORY_POLICY - the eviction policy the valkey service uses once VALKEY_MAXMEMORY is reached (default allkeys-lru)VALKEY_AUTO_AOF_REWRITE_MIN_SIZE - the minimum size the append-only file must reach before an automatic AOF rewrite is triggered (default 64mb)valkey-cache service:
VALKEY_CACHE_MAXMEMORY, VALKEY_CACHE_MAXMEMORY_POLICY - the valkey-cache service’s equivalents of VALKEY_MAXMEMORY/VALKEY_MAXMEMORY_POLICY above; note that valkey-cache runs with appendonly no, i.e. without AOF persistence, since it’s a cache rather than a durable storezeek.env, zeek-live.env and zeek-offline.env - settings for Zeek and for scanning extracted files Zeek observes in network traffic
ZEEK_AUTO_ANALYZE_PCAP_FILES - if set to true, all PCAP files imported into Malcolm will automatically be analyzed by Zeek, and the resulting logs will also be imported (default false)ZEEK_AUTO_ANALYZE_PCAP_THREADS - the number of threads available to Malcolm for analyzing Zeek logs (default 1)ZEEK_ZAM - if set to true, enables Zeek’s ZAM script optimizer (-O ZAM)ZEEK_JSON - whether Zeek should generate JSON format logs (true) or TSV format logs (false)ZEEK_…_DETAILED - some network analyzers produce can produce more detailed records in addition to “summary” records; set these variables to true to turn on the more verbose versionZEEK_DISABLE_… - if set to true, each of these variables can be used to disable a certain Zeek function or network analyzer when it analyzes PCAP files (for example, setting ZEEK_DISABLE_LOG_PASSWORDS to true to disable logging of cleartext passwords)ZEEK_DISABLE_SPICY_ZIP - with Strelka’s archive scanners enabled, this should be true to avoid duplicate processingZEEK_…_PORTS - used to specify non-default ports to register certain Zeek analyzers (e.g., ZEEK_SYNCHROPHASOR_PORTS for the ICSNPP-Synchrophasor analyzer, ZEEK_GENISYS_PORTS for the ICSNPP-Genisys analyzer, and ZEEK_ENIP_PORTS for the ICSNPP-Ethernet/IP analyzer) formatted as a comma-separated list of Zeek ports (e.g., 12345/tcp or 4041/tcp,4042/udp)ZEEK_DISABLE_INTEL_OFFLINE/ZEEK_DISABLE_INTEL_LIVE - if set to true, Zeek will not @load the files under ./zeek/intel as described in Zeek Intelligence Framework for historical PCAP processing/live traffic capture, respectivelyZEEK_DISABLE_ICS_ALL and ZEEK_DISABLE_ICS_… - if set to true, these variables can be used to disable Zeek’s protocol analyzers for Operational Technology/Industrial Control Systems (OT/ICS) protocolsZEEK_DISABLE_BEST_GUESS_ICS - see “Best Guess” Fingerprinting for ICS ProtocolsZEEK_EXTRACTOR_MODE - determines the file extraction behavior for file transfers detected by Zeek; see Automatic file extraction and scanning for more detailsEXTRACTED_FILE_MAX_BYTES - the maximum size (in bytes) for files to be extracted by ZeekZEEK_FILE_ANALYZER_TIMEOUT_SEC - default amount of time a file can be inactive before the file analysis gives up and discards any internal state related to the fileZEEK_INTEL_FEED_SINCE - when querying a TAXII, MISP, Google, or Mandiant threat intelligence feed, only process threat indicators created or modified since the time represented by this value; it may be either a fixed date/time (01/01/2025) or relative interval (24 hours ago). Note that this value can be overridden per-feed by adding a since: value to each feed’s respective configuration YAML file.ZEEK_INTEL_FEED_SSL_CERTIFICATE_VERIFICATION - whether or not to require SSL certificate verification when querying an intelligence feed (default false)ZEEK_INTEL_REFRESH_THREADS - number of threads to use for querying feeds for generating Zeek Intelligence Framework filesZEEK_INTEL_ITEM_EXPIRATION - specifies the value for Zeek’s Intel::item_expiration timeout as used by the Zeek Intelligence Framework (default -1min, which disables item expiration)ZEEK_INTEL_REFRESH_CRON_EXPRESSION - specifies a cron expression (using cronexpr-compatible syntax) indicating the refresh interval for generating the Zeek Intelligence Framework files (defaults to empty, which disables automatic refresh)ZEEK_INTEL_REFRESH_ON_STARTUP - if set to true, Zeek intelligence framework files will be refreshed upon startupZEEK_JA4SSH_PACKET_COUNT - the Zeek JA4+ plugin calculates the JA4SSH value once for every x SSH packets; x is set here (default 200)ZEEK_KAFKA_ENABLED - if set to true, Zeek will send all active logs to Kafka via the zeek-kafka plugin, in JSON formatZEEK_KAFKA_BROKERS - the Kafka broker list to send logs to (Kafka’s metadata.broker.list)ZEEK_KAFKA_TOPIC - the Kafka topic name to which Zeek logs are publishedZEEK_LONG_CONN_DURATIONS - a comma-separated list of durations, in seconds, at which point “long connections” will be logged (default 300,600,1800,3600,43200,86400)ZEEK_LONG_CONN_DO_NOTICE - if set to true, a notice.log entry will be created when the zeek-long-connections plugin discovers what it considers to be a long connection (default true)ZEEK_LONG_CONN_REPEAT_LAST_DURATION - if set to true, logging will be repeated at the last interval specified in ZEEK_LONG_CONN_DURATIONS (default true)ZEEK_LIVE_CAPTURE - if set to true, Zeek will monitor live traffic on the local interface(s) defined by PCAP_FILTER
ZEEK_DISABLE_STATS - if ZEEK_LIVE_CAPTURE is true and this variable is set to false or blank, Malcolm will enable capture statistics Zeek, which data is used to populate the Packet Capture Statistics dashboardZEEK_LOCAL_NETS - specifies the value for Zeek’s Site::local_nets variable (and networks.cfg for live capture) (e.g., 1.2.3.0/24,5.6.7.0/24); note that by default, Zeek considers IANA-registered private address space such as 10.0.0.0/8 and 192.168.0.0/16 site-localZEEK_ROTATED_PCAP - if set to true, Zeek can analyze captured PCAP files captured by netsniff-ng or tcpdump (see PCAP_ENABLE_NETSNIFF and PCAP_ENABLE_TCPDUMP, as well as ZEEK_AUTO_ANALYZE_PCAP_FILES); if ZEEK_LIVE_CAPTURE is true, this should be false; otherwise Zeek will see duplicate trafficThe ./scripts/configure script can also be run noninteractively which can be useful for scripting Malcolm setup. This behavior can be selected by supplying the -d or --defaults option on the command line. Running with the --help option will list the arguments accepted by the script:
usage: configure [-h] [--debug [true|false]] [--quiet] [--configure [true|false]] [--dry-run] [--log-to-file [filename]] [--skip-splash] [--tui | --dui | --gui | --non-interactive] [--compose-file <string>] [--environment-dir-input <string>] [--environment-dir-output <string>]
[--export-malcolm-config-file [<path>]] [--import-malcolm-config-file <path> | --load-existing-env [true|false] | --defaults] [--malcolm-file <string>] [--image-file <string>] [--extra [EXTRASETTINGS ...]]
Malcolm Installer
options:
-h, --help show this help message and exit
Installer Options:
--debug, --verbose [true|false]
Enable debug output including tracebacks and debug utilities
--quiet, --silent Suppress console logging output during installation
--configure, -c [true|false]
Only write configuration and ancillary files; skip installation steps
--dry-run Log planned actions without writing files or making system changes
--log-to-file [filename]
Log output to file. If no filename provided, creates timestamped log file.
--skip-splash Skip the splash screen prompt on startup
Interface Mode (mutually exclusive):
--tui Run in command-line text-based interface mode (default)
--dui Run in python dialogs text-based user interface mode (if available - requires python dialogs)
--gui Run in graphical user interface mode (if available - requires customtkinter)
--non-interactive Run in non-interactive mode for unattended installations (suppresses all user prompts)
Configuration File Options:
--compose-file, --configure-file, --kube-file, -f <string>
Path to docker-compose.yml (for compose) or kubeconfig (for Kubernetes)
Environment Config Options:
--environment-dir-input <string>
Input directory containing Malcolm's .env and .env.example files
--environment-dir-output, -e <string>
Target directory for writing Malcolm's .env files
--export-malcolm-config-file, --export-mc-file [<path>]
Export configuration to JSON/YAML settings file (auto-generates filename if not specified)
--import-malcolm-config-file, --import-mc-file <path>
Import configuration from JSON/YAML settings file
--load-existing-env, -l [true|false]
Automatically load provided config/ .env files from the input directory when present. Can be used in conjunction with --environment-dir-input
--defaults, -d Use built-in default configuration values and skip loading from the config directory
Installation Files:
--malcolm-file, -m <string>
Malcolm .tar.gz file for installation
--image-file, -i <string>
Malcolm container images .tar.xz file for installation
Additional Configuration Options:
--extra [EXTRASETTINGS ...]
Extra environment variables to set (e.g., foobar.env:VARIABLE_NAME=value)
…
Once Malcolm is configured correctly, the --export-malcolm-config-file option can be used to export the configuration to a file that can be used with --import-malcolm-config-file to restore it later or transfer it to another Malcolm instance for import.
To modify Malcolm settings programmatically in scripting, a tool like jq can be used with --export-malcolm-config-file and --import-malcolm-config-file, as illustrated here:
# export the current configuration to a JSON file without modifying anything in ./config/
SETTINGS_FILE="$(mktemp --suffix=.json)"
./scripts/configure --dry-run --non-interactive --export-malcolm-config-file "${SETTINGS_FILE}"
# use JQ To set whatever options in the exported JSON configuration file you wish to change
JQ_FILE="$(mktemp --suffix=.jq)"
tee "${JQ_FILE}" >/dev/null <<EOF
.configuration.dashboardsDarkMode = true
| .configuration.reverseDns = true
| .configuration.pcapNodeName = "Engineering Workstation"
EOF
jq -f "${JQ_FILE}" "${SETTINGS_FILE}" | sponge "${SETTINGS_FILE}"
# import the modified configuration
./scripts/configure --non-interactive --import-malcolm-config-file "${SETTINGS_FILE}"
# clean up
rm -f "${SETTINGS_FILE}" "${JQ_FILE}"
Similarly, authentication-related settings can also be set noninteractively by using the command-line arguments for ./scripts/auth_setup.
In instances where Malcolm is deployed with the intention of running indefinitely, eventually the question arises of what to do when the file systems used for storing Malcolm’s artifacts (e.g., PCAP files, raw logs, OpenSearch indices, extracted files, etc.). Malcolm provides options for tuning the “aging out” (deletion) of old artifacts to make room for newer data.
arkime.env:
MANAGE_PCAP_FILES - if set to true, all PCAP files imported into Malcolm will be marked as available for deletion by Arkime if available storage space becomes too low (default false)ARKIME_FREESPACEG - when MANAGE_PCAP_FILES is true, this value is used by Arkime to determine when to delete the oldest PCAP files. Note that this variable represents the amount of free/unused/available desired on the file system: e.g., a value of 5% means “delete PCAP files if the amount of unused storage on the file system falls below 5%” (default 10%).filebeat.env:
LOG_CLEANUP_MINUTES - specifies the age, in minutes, at which already-processed log files should be deletedZIP_CLEANUP_MINUTES - specifies the age, in minutes, at which the compressed archives containing already-processed log files should be deleted./zeek-logs/extract_files/ directory can be periodically pruned based on the following variables in zeek.env. If either of the two threshold limits defined here are met, the oldest extracted files will be deleted until the limit is no longer met. Setting either of the threshold limits to 0 disables that check.
FILESCAN_PRUNE_THRESHOLD_MAX_SIZE - specifies the maximum size, specified either in gigabytes or as a human-readable data size (e.g., 250G), that the ./zeek-logs/extract_files/ directory is allowed to contain before the prune condition triggersFILESCAN_PRUNE_THRESHOLD_TOTAL_DISK_USAGE_PERCENT - specifies a maximum fill percentage for the file system containing the ./zeek-logs/extract_files/; in other words, if the disk is more than this percentage utilized, the prune condition triggersFILESCAN_PRUNE_INTERVAL_SECONDS - the interval between checking the prune conditions, in seconds (default 300)OPENSEARCH_INDEX_SIZE_PRUNE_LIMIT variable in dashboards-helper.env defines a maximum cumulative that OpenSearch indices are allowed to consume before the oldest indices are deleted, specified as either as a human-readable data size (e.g., 250G) or as a percentage of the total disk size (e.g., 70%): e.g., a value of 500G means “delete the oldest OpenSearch indices if the total space consumed by Malcolm’s indices exceeds five hundred gigabytes.”