LILT is building a rigorous evaluation suite for large language models to test multilingual robustness in terminal environments. This role involves designing and validating benchmark tasks to measure AI performance on complex, non-English software challenges.
Remote (Rio de Janeiro, Rio de Janeiro)
$50 - $75/hr
Contract
Entry Level
1 perks
Posted 1 day ago
About the role
LILT is building a rigorous evaluation suite for large language models to test multilingual software challenges. This role involves designing and validating benchmark tasks that measure model robustness in non-English terminal workflows.
Ignite IT is seeking a Junior AI Prompt Engineer to support U.S. Customs and Border Protection by identifying AI adoption opportunities, designing prompts, and developing AI-enabled workflows to improve operational efficiency and mission outcomes.
LILT is seeking a native French-speaking software engineer to design and validate multilingual benchmarks for large language models. This role focuses on creating rigorous evaluation tasks that test AI robustness in non-English terminal workflows and locale-specific edge cases.
Skills
PythonShell ScriptingData ProcessingTerminalCLICoding AgentsUnicodeMultilingual Text ProcessingSoftware EngineeringQuality AssuranceBenchmark EngineeringPromptingTranslationDebuggingTechnical Writing
LILT is seeking a Native Language Specialist in French (Canada) to design and validate rigorous multilingual benchmarks for large language models. This role focuses on creating high-quality, verifiable tasks that test AI robustness in complex, non-English terminal workflows.
Skills
PythonShell ScriptingData ProcessingTerminalCLICoding AgentsUnicodeMultilingual Text ProcessingSoftware EngineeringQuality AssurancePromptingBenchmark DevelopmentDebuggingTechnical Writing
LILT is seeking a Native French-speaking AI Benchmark Engineer to design and validate rigorous multilingual software tasks for evaluating large language models. This role involves creating realistic task environments and deterministic verifier scripts to measure model robustness in non-English terminal workflows.
LILT is building a rigorous evaluation suite for large language models to test multilingual robustness in terminal workflows. This role involves designing and validating high-quality benchmark tasks for coding agents using native Spanish (Argentina) expertise.
LILT is building a rigorous evaluation suite for large language models to test multilingual software challenges. This role involves designing and validating terminal-based benchmarks to measure model robustness in non-English environments.
LILT is building a rigorous evaluation suite for large language models to test multilingual software challenges. This role involves designing and validating high-quality benchmark tasks to measure model robustness in non-English environments.
LILT is building a rigorous evaluation suite for large language models to test multilingual software challenges. This role involves designing and validating terminal-based benchmarks to measure model robustness in non-English environments.
Rank Interactive, a leading global digital gaming company, is seeking an Agentic AI Designer to shape autonomous, multi-step AI-powered customer journeys. This role bridges technology and design to optimize customer and agent interactions, ensuring they are efficient, human-like, and compliant with regulatory standards.
State Street is a global investment management and custodian bank seeking a Prompting Execution Lead to scale Agentic AI capabilities. The role focuses on operationalizing prompting frameworks, coordinating testing, and integrating AI solutions into business-as-usual operations.
Pavago is a staffing and recruiting firm seeking an AI Operations Specialist to manage and optimize production AI conversation systems for telehealth clients. The role involves configuring workflows, monitoring performance, and refining prompts to improve client conversion rates.
Klipboard (formerly Kerridge Commercial Systems) is a global provider of integrated ERP and business management software for the distributive trade sector. This senior engineering role focuses on scaling AI practices by designing prompt strategies, establishing evaluation standards, and mentoring engineers to embed reliable, measurable AI features into production products.
AstraZeneca is seeking an AI Prompt & Market Rollout Specialist to orchestrate the global deployment of an AI Coach platform across 40–50 markets. This role focuses on prompt engineering, avatar adaptation, and stakeholder coordination to enhance field engagement and commercial excellence.
Hybrid (Chortiatis Municipal Unit, Macedonia and Thrace)
Full-time
Mid Level
1 perks
Posted 2 days ago
About the role
Pfizer is seeking a Manager, AI Ops Site Reliability Engineer to design and evaluate AI agent workflows for operational incident management. This role focuses on automating operations by building evaluation harnesses and converting domain expertise into reliable, measurable agent behavior.
Skills
AI OpsSite Reliability EngineeringPythonLLMPrompt EngineeringAgent FrameworksRAGMachine LearningData EngineeringObservabilityServiceNowDynatraceAgileInfrastructure OperationsCloud Services
Prophetic is building an AI-native platform for land acquisition in real estate, helping homebuilders and developers analyze opportunities faster. This role owns the ground truth and evaluation datasets to ensure high-quality model performance across the platform.
Skills
Machine learningLLM evaluationPythonSQLData labelingGround truth developmentStatistical analysisFeature engineeringPrompt engineeringAgentic workflowsModel validationData lineagePrecision and recallA/B testingSystem architecture
MPOWERHealth is a healthcare services company focused on surgical assist and intraoperative neuromonitoring. This role leads the practical adoption of enterprise generative AI by partnering with department leaders to develop secure, repeatable AI-assisted workflows and ensure compliance with privacy and regulatory standards.
Ciklum is a global engineering firm seeking a Senior AI Evaluation Engineer to build central evaluation harnesses and golden datasets for cross-functional teams. The role focuses on setting pass thresholds, implementing production monitoring, and ensuring safety and medical accuracy in AI systems.
Citco, a global leader in asset servicing, seeks a Conversation Designer to build and govern the knowledge base for its AI assistant systems. The role focuses on structuring content, designing prompt behaviors, and ensuring regulatory compliance to improve AI accuracy and user experience.
LILT is building a rigorous evaluation suite for large language models to test multilingual robustness in terminal environments. This role involves designing and validating benchmark tasks to measure AI performance on complex, non-English software challenges.
Remote (Rio de Janeiro, Rio de Janeiro)
$50 - $75/hr
Contract
Entry Level
1 perks
Posted 1 day ago
About the role
LILT is building a rigorous evaluation suite for large language models to test multilingual software challenges. This role involves designing and validating benchmark tasks that measure model robustness in non-English terminal workflows.
Ignite IT is seeking a Junior AI Prompt Engineer to support U.S. Customs and Border Protection by identifying AI adoption opportunities, designing prompts, and developing AI-enabled workflows to improve operational efficiency and mission outcomes.
LILT is seeking a native French-speaking software engineer to design and validate multilingual benchmarks for large language models. This role focuses on creating rigorous evaluation tasks that test AI robustness in non-English terminal workflows and locale-specific edge cases.
Skills
PythonShell ScriptingData ProcessingTerminalCLICoding AgentsUnicodeMultilingual Text ProcessingSoftware EngineeringQuality AssuranceBenchmark EngineeringPromptingTranslationDebuggingTechnical Writing
LILT is seeking a Native Language Specialist in French (Canada) to design and validate rigorous multilingual benchmarks for large language models. This role focuses on creating high-quality, verifiable tasks that test AI robustness in complex, non-English terminal workflows.
Skills
PythonShell ScriptingData ProcessingTerminalCLICoding AgentsUnicodeMultilingual Text ProcessingSoftware EngineeringQuality AssurancePromptingBenchmark DevelopmentDebuggingTechnical Writing
LILT is seeking a Native French-speaking AI Benchmark Engineer to design and validate rigorous multilingual software tasks for evaluating large language models. This role involves creating realistic task environments and deterministic verifier scripts to measure model robustness in non-English terminal workflows.
LILT is building a rigorous evaluation suite for large language models to test multilingual robustness in terminal workflows. This role involves designing and validating high-quality benchmark tasks for coding agents using native Spanish (Argentina) expertise.
LILT is building a rigorous evaluation suite for large language models to test multilingual software challenges. This role involves designing and validating terminal-based benchmarks to measure model robustness in non-English environments.
LILT is building a rigorous evaluation suite for large language models to test multilingual software challenges. This role involves designing and validating high-quality benchmark tasks to measure model robustness in non-English environments.
LILT is building a rigorous evaluation suite for large language models to test multilingual software challenges. This role involves designing and validating terminal-based benchmarks to measure model robustness in non-English environments.
Rank Interactive, a leading global digital gaming company, is seeking an Agentic AI Designer to shape autonomous, multi-step AI-powered customer journeys. This role bridges technology and design to optimize customer and agent interactions, ensuring they are efficient, human-like, and compliant with regulatory standards.
State Street is a global investment management and custodian bank seeking a Prompting Execution Lead to scale Agentic AI capabilities. The role focuses on operationalizing prompting frameworks, coordinating testing, and integrating AI solutions into business-as-usual operations.
Pavago is a staffing and recruiting firm seeking an AI Operations Specialist to manage and optimize production AI conversation systems for telehealth clients. The role involves configuring workflows, monitoring performance, and refining prompts to improve client conversion rates.
Klipboard (formerly Kerridge Commercial Systems) is a global provider of integrated ERP and business management software for the distributive trade sector. This senior engineering role focuses on scaling AI practices by designing prompt strategies, establishing evaluation standards, and mentoring engineers to embed reliable, measurable AI features into production products.
AstraZeneca is seeking an AI Prompt & Market Rollout Specialist to orchestrate the global deployment of an AI Coach platform across 40–50 markets. This role focuses on prompt engineering, avatar adaptation, and stakeholder coordination to enhance field engagement and commercial excellence.
Hybrid (Chortiatis Municipal Unit, Macedonia and Thrace)
Full-time
Mid Level
1 perks
Posted 2 days ago
About the role
Pfizer is seeking a Manager, AI Ops Site Reliability Engineer to design and evaluate AI agent workflows for operational incident management. This role focuses on automating operations by building evaluation harnesses and converting domain expertise into reliable, measurable agent behavior.
Skills
AI OpsSite Reliability EngineeringPythonLLMPrompt EngineeringAgent FrameworksRAGMachine LearningData EngineeringObservabilityServiceNowDynatraceAgileInfrastructure OperationsCloud Services
Prophetic is building an AI-native platform for land acquisition in real estate, helping homebuilders and developers analyze opportunities faster. This role owns the ground truth and evaluation datasets to ensure high-quality model performance across the platform.
Skills
Machine learningLLM evaluationPythonSQLData labelingGround truth developmentStatistical analysisFeature engineeringPrompt engineeringAgentic workflowsModel validationData lineagePrecision and recallA/B testingSystem architecture
MPOWERHealth is a healthcare services company focused on surgical assist and intraoperative neuromonitoring. This role leads the practical adoption of enterprise generative AI by partnering with department leaders to develop secure, repeatable AI-assisted workflows and ensure compliance with privacy and regulatory standards.
Ciklum is a global engineering firm seeking a Senior AI Evaluation Engineer to build central evaluation harnesses and golden datasets for cross-functional teams. The role focuses on setting pass thresholds, implementing production monitoring, and ensuring safety and medical accuracy in AI systems.
Citco, a global leader in asset servicing, seeks a Conversation Designer to build and govern the knowledge base for its AI assistant systems. The role focuses on structuring content, designing prompt behaviors, and ensuring regulatory compliance to improve AI accuracy and user experience.